System
The system addresses the lack of conversational opportunities for language learners by providing high-quality, real-time interactions with specific experts and celebrities, while ensuring efficient data processing and fair reward distribution for providers.
Patent Information
- Application Number
- JP2024116573
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-19
- Publication Date
- 2026-01-29
AI Technical Summary
Language learners lack opportunities for real-life conversational experiences and direct interaction with experts or celebrities, limiting their ability to improve language skills effectively, while existing systems have issues with data quality, model specificity, and inefficient reward distribution.
A system that collects, cleans, and stores conversation data in a database, generates a large-scale language model tailored to specific partners, provides real-time conversations, and distributes rewards to data providers based on usage.
Enhances language learning through realistic conversational experiences and provides data providers with new revenue streams by ensuring high-quality data processing and transparent reward mechanisms.
Smart Images

Figure 2026015099000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Describe the "problem that the invention aims to solve" and the "means for solving the problem."
[0005] With the advancement of globalization and technology, the ability to understand different languages and cultures is becoming increasingly important. However, language learners have limited opportunities to improve their skills through real-life conversational experiences. Furthermore, direct interaction with experts or celebrities is generally difficult. Therefore, new methods are needed to help language learners stay motivated and progress effectively. [Means for solving the problem]
[0006] The present invention includes a means for collecting, cleaning, and storing conversation data in a database. It also includes a means for generating a large-scale language model for training using the stored conversation data, and for providing an appropriate AI model based on a specific conversation partner after a user is authenticated. It also includes a means for conducting a real-time conversation between the user and the AI model, and for the AI model to generate responses based on the user's input. This allows language learners to effectively improve their skills through realistic conversational experiences. It also includes a means for calculating and distributing rewards to data providers based on the conversation data used, and the system offers an attractive reward model for conversation data providers.
[0007] Understood. Below are definitions of important terms included in the claims.
[0008] "Conversation Data" refers to dialogue information in the form of audio, text, or video provided by a person with expert knowledge or a celebrity.
[0009] "Cleaning" is the process of removing noise and unnecessary parts from collected conversation data and shaping it into an appropriate format.
[0010] A "database" is a system for systematically storing cleaned conversation data and making it easily accessible.
[0011] A "large-scale language model" is an AI model that performs natural language processing based on large amounts of text data, and uses conversational data to generate responses that are tailored to specific individuals.
[0012] A "user" is a language learner who engages in conversation simulations with a specific conversation partner through the system.
[0013] "Authentication" is the identity verification process required for a user to access a system.
[0014] A "conversation partner" is a celebrity or professional person that the user simulates.
[0015] "AI Model" means an artificial intelligence program that uses a large-scale language model to generate responses for a specific conversation partner.
[0016] "Real-time" means that responses are generated immediately in response to user input, and the conversation proceeds without delay.
[0017] "Reward" is monetary compensation paid to the conversation data provider based on the conversation data used. [Brief explanation of the drawings]
[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0020] First, the terms used in the following description will be explained.
[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0026] [First embodiment]
[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0039] The present invention provides a system that collects, cleans, and stores conversation data in a database, generates a large-scale language model for training using the stored conversation data, and provides an appropriate AI model based on a specific conversation partner after a user is authenticated. The system also conducts real-time conversations between the user and the AI model, and the AI model generates responses based on the user's input. The system also includes a means for calculating and distributing rewards to data providers based on the conversation data used.
[0040] Program processing
[0041] 1. Collecting conversation data
[0042] Server: Provides an interface for conversation data providers (celebrities and professionals) to upload their conversation data after logging in. When providers upload audio, text, or video data, the data is temporarily stored.
[0043] 2. Data Preprocessing
[0044] Server: Analyzes the received conversation data, removes noise and unnecessary parts, and filters out personal information. The preprocessed data is saved in the database as cleaned data.
[0045] 3. AI model generation
[0046] Server: Uses the cleaned data to train large-scale language models, which are then tuned to specific conversational partners and set up API endpoints to make them accessible.
[0047] 4. User Authentication and Selection
[0048] Terminal: Provides an authentication interface for users to access the system through the application and log in. Once authenticated, users can select a specific conversation partner from a list.
[0049] 5. Conversation Simulation
[0050] Terminal: The terminal displays an input form on the interface to start a conversation with a conversation partner of the user's choice. The user enters the information and sends it to the server.
[0051] Server: Receives user input and passes it to the appropriate AI model, which generates a response and sends it back to the device.
[0052] Terminal: Displays responses from the server and provides an environment in which the user can continue the dialogue.
[0053] 6. Reward Distribution
[0054] Server: Tracks the user's conversation time and frequency of use, and calculates the reward for the conversation data provider. Once the reward is calculated, payment is processed based on the provider's account information.
[0055] Specific examples
[0056] For example, consider the case where a famous actor B provides audio data about his or her stage experience. This data is collected, cleaned to remove personal information, and then stored in a database. Next, a large-scale language model is trained based on Actor B's audio data to generate an AI model that can converse with Actor B.
[0057] A user logs into this system and selects "Simulated dialogue with actor B." When the user inputs "What was your most recent stage challenge?" to actor B, the input is sent to the server, and the AI model generates a response such as "My most recent stage challenge was trying out a new role," and sends it back to the user.
[0058] When the conversation ends, the server calculates the reward for Actor B based on the time and frequency of use and transfers it to the system. In this way, users can advance their language learning through a realistic dialogue experience, and the provider of the conversation data can also earn rewards.
[0059] This system not only allows language learners to stay motivated and improve their skills effectively, but also allows conversation data providers to earn revenue in new ways.
[0060] The processing flow will be explained below.
[0061] Understood. Below is the process flow broken down into specific steps.
[0062] Step 1:
[0063] Server: Provides an authentication interface for conversation data providers (e.g., famous actor B) to log in. The conversation data provider logs in by entering the correct authentication information.
[0064] Step 2:
[0065] Server: After successful authentication, the server displays a conversation data upload interface to the conversation data provider, where the provider can select and upload conversation data in audio, text, or video format.
[0066] Step 3:
[0067] Server: Temporarily stores the uploaded data, checks the format and content of the data, and if the format is correct, sends the data to the cleaning process.
[0068] Step 4:
[0069] Server: The cleaning process removes noise and unnecessary parts from the conversation data and filters out personal information, such as the patient's name and contact details.
[0070] Step 5:
[0071] Server: Organizes the cleaned data and stores it in a database. When storing the data, it categorizes it by conversation topic, person, date and time, etc.
[0072] Step 6:
[0073] Server: Uses the cleaned conversation data to train a large-scale language model (LLM). For example, it trains the model based on actor B's data and builds a response generation model specialized for that person.
[0074] Step 7:
[0075] Server: Evaluates the trained model on the test data, adjusts the model parameters if necessary, saves the completed model, and sets up the appropriate API endpoints.
[0076] Step 8:
[0077] Terminal: The user launches the application and accesses the login screen. The user enters their credentials and logs in.
[0078] Step 9:
[0079] Server: Checks the user's authentication information, and if authentication is successful, displays the user dashboard, where the user can select the conversation partner they want to simulate (e.g., Actor B).
[0080] Step 10:
[0081] Terminal: Displays an interface for the user to select a simulated conversation partner. The user makes the selection and sends the selection information to the server.
[0082] Step 11:
[0083] Server: Based on the user's selection, prepares the corresponding AI model (e.g., actor B's model) and sends interface information to the terminal to start the conversation.
[0084] Step 12:
[0085] Terminal: Sends a text message typed by the user to the server. For example, "What's your latest stage challenge?"
[0086] Step 13:
[0087] Server: Receives user input and passes it to the appropriate AI model, which then generates a response, for example, "My most recent stage challenge was taking on a new role."
[0088] Step 14:
[0089] Server: Sends the generated response to the user's device.
[0090] Step 15:
[0091] Terminal: Receives the response from the server and displays it in the user interface. If the user wants to continue the interaction, they can input again.
[0092] Step 16:
[0093] Server: After the user's simulation is completed, the server calculates the reward for the conversation data provider based on the simulation time and number of times used.
[0094] Step 17:
[0095] Server: Carries out the payment procedures to distribute the calculated reward to the conversation data provider. The reward is transferred based on the provider's account information.
[0096] The above is the specific processing flow of this system.
[0097] Example 1
[0098] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0099] Conventional conversation data collection and processing systems have insufficient quality control of collected data, particularly due to the inclusion of noise and personal information, which limits its use. Furthermore, when training AI models using conversation data, it is difficult to adjust them to specific conversation partners, making it difficult to provide real-time responses to users. Furthermore, reward distribution to data providers is sometimes not transparent and efficient, resulting in a lack of incentives for providers.
[0100] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0101] In this invention, the server includes a means for collecting conversation data, a means for cleaning the collected conversation data and storing it in a database, and a means for generating a model to be trained using the stored conversation data. This enables appropriate preprocessing of the data and the generation of an AI model tailored to a specific conversation partner using high-quality training data. The server also includes a means for selecting a specific conversation partner after the user is authenticated, a means for conducting a conversation between the user and the model and for the model to generate a response based on the user's input, and a means for calculating and distributing rewards to data providers based on the used conversation data. This allows users to receive appropriate responses in real time and distributes rewards fairly and quickly to data providers, improving the efficiency and satisfaction of the entire system.
[0102] "Conversation data" refers to audio data, text data, or video data provided to the system by users or conversation data providers.
[0103] "Cleaning" is the process of removing noise, inappropriate content, personal information, etc. from collected conversation data to improve data quality before storing it in a database.
[0104] A "database" is a storage within the system that efficiently and securely stores cleaned conversation data for later use as training data.
[0105] "Training" is the process of using cleaned conversational data to train an AI model and improve its pattern recognition capabilities and prediction accuracy.
[0106] A "model" is a collection of AI structures and algorithms trained on conversational data that are used to generate appropriate responses to user input.
[0107] A "specific conversation partner" is a virtual conversation partner for a dialogue that is selected by the user based on various categories and profiles provided on the system.
[0108] "Real-time" means that there is a very short delay between when a user makes an input and when the AI model returns a response, and refers to a situation in which the interaction takes place immediately.
[0109] A "response" is a response, such as text or voice, that an AI model generates based on user input.
[0110] "Remuneration" is the compensation calculated and paid to the conversation data provider based on the degree and frequency of use of the provided data.
[0111] "Distribution" is the process of appropriately paying the calculated reward based on the registered information of the conversation data provider.
[0112] "Interface" refers to the screens, input forms, and operating means used by conversation data providers when uploading data or by users when accessing and operating the system.
[0113] To implement this invention, the server, terminal, and user each play specific roles. The main hardware used includes a database server, web server, and user terminal (PC or smartphone), while the software required includes a framework for training large-scale language models (e.g., TensorFlow, PyTorch), database software (e.g., MySQL, PostgreSQL), and a web service API.
[0114] The server provides a means for collecting conversation data. Specifically, it provides a user interface for conversation data providers to log in and upload audio data, text data, and video data. The device supports a means for providers to upload data, allowing them to select and send data in various formats (e.g., .mp3, .txt, .mp4). Uploaded data is temporarily stored on the server.
[0115] The server analyzes the collected conversation data, removes noise and unnecessary parts, and filters out personal information to clean it. The cleaned data is then stored in a database, which later serves as the foundation for training AI models.
[0116] The server then uses the cleaned conversation data to train a large-scale language model using a large-scale language model framework, processing the data in batches to learn the model weights, resulting in an AI model that corresponds to a specific conversation partner. The model is then made accessible through an API endpoint.
[0117] To access the system, a user uses a terminal to enter a username and password into a login interface. After successful authentication, the server provides the user with a list of specific available conversation partners from which the user can select a conversation partner.
[0118] The device then displays an interface for the user to converse with the selected conversation partner. The user enters a question or message into the input form and clicks the send button. For example, a question such as "What was your most recent stage challenge?" is sent to the server.
[0119] The server receives user input, passes it to the trained AI model, and generates a response, which is then sent back to the device, where it is displayed to the user, allowing the user to continuously engage in real-time dialogue with the AI model.
[0120] Finally, the server tracks the duration and frequency of users' conversation data usage and calculates rewards for conversation data providers based on that data. The calculated rewards are distributed based on the providers' registered account information.
[0121] As a concrete example, consider a case where a user selects a dialogue simulation with actor B and asks, "What was the most moving moment on stage?" This question is transmitted to the AI model via the server, and the AI model generates a response such as, "The most moving moment on stage was when the audience shed tears," and sends it back to the user.
[0122] This system allows users to advance their language learning through realistic conversational experiences, and enables conversation data providers to earn revenue in new ways.
[0123] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0124] Step 1:
[0125] The server verifies the login of the conversation data provider and, after successful authentication, provides an interface for data upload.
[0126] Input: Provider username and password.
[0127] Specific operation: Query the database to verify the provider's authentication information, and if authentication is successful, start a login session.
[0128] Output: Login session token and data upload interface.
[0129] Step 2:
[0130] Using the terminal, the conversation data provider selects audio data, text data, or video data and clicks the upload button.
[0131] Input: Data files selected by the provider (e.g. .mp3, .txt, .mp4).
[0132] Specific operation: A file selection dialog is displayed, and after the provider selects a file, they press the upload button.
[0133] Output: The selected data file is sent to the server.
[0134] Step 3:
[0135] The server stores the uploaded data in temporary storage and checks the integrity of the data.
[0136] Input: Uploaded data file.
[0137] What it does: Saves the file to temporary storage and verifies the data integrity by checking the file format and the first few bytes of the data.
[0138] Output: Data file with integrity checked.
[0139] Step 4:
[0140] The server analyzes the received conversation data, removes noise and unnecessary parts, and filters out personal information to clean it.
[0141] Input: Data file whose integrity has been checked.
[0142] Specific operation: For audio data, a noise reduction algorithm is applied, and for text data, filtering is performed to mask specific keywords.
[0143] Output: Cleaned data.
[0144] Step 5:
[0145] The server organizes and stores the cleaned data in a database.
[0146] Input: Cleaned data.
[0147] What it does: Inserts data and creates indexes according to the appropriate schema in the database.
[0148] Output: Cleaned data stored in a database.
[0149] Step 6:
[0150] The server uses the cleaned conversational data to train a large-scale language model.
[0151] Input: Cleaned data in the database.
[0152] Specific behavior: Use a large-scale language model framework (e.g., TensorFlow, PyTorch) to process data in batches and update the model weights.
[0153] Output: A trained AI model.
[0154] Step 7:
[0155] The server tunes the trained model to correspond to specific conversation partners and sets up an API endpoint to make it accessible externally.
[0156] Input: A trained AI model.
[0157] Specific actions: Adjust model parameters, configure API endpoints, and deploy.
[0158] Output: An accessible API endpoint for the AI model.
[0159] Step 8:
[0160] The device presents the user with a login interface, prompting them to enter their username and password, and upon successful authentication, providing an interface for selecting a specific conversation partner.
[0161] Input: Username and Password.
[0162] Specific behavior: The user enters information into the login form and clicks the login button.
[0163] Output: Conversation partner selection interface for authenticated user.
[0164] Step 9:
[0165] The terminal displays a screen for starting a conversation based on the conversation partner selected by the user.
[0166] Input: User's selected conversation partner information.
[0167] Specific behavior: Displays a conversation interface and provides an input form.
[0168] Output: A form for the user to enter.
[0169] Step 10:
[0170] The user enters a question or message into the input form and clicks the send button.
[0171] Input: User question or message (e.g., "What was your most recent stage challenge?").
[0172] Specific behavior: The user enters text and clicks the submit button.
[0173] Output: The entered question or message is sent to the server.
[0174] Step 11:
[0175] The server receives user input and passes it to a trained AI model to generate a response.
[0176] Input: The user's typed question or message.
[0177] What it does: Tokenizes the user's message and passes it to the AI model for processing.
[0178] Output: The generated response of the AI model.
[0179] Step 12:
[0180] The terminal displays responses from the server to the user, providing an interface that allows the conversation to continue.
[0181] Input: The response sent back from the server.
[0182] Specific behavior: Displays the response text and persists the interface so the user can enter it again.
[0183] Output: The response displayed to the user and an input form for further interaction.
[0184] Step 13:
[0185] The server tracks the time and frequency of users' conversation data usage and calculates rewards for conversation data providers based on that data.
[0186] Input: User conversation time, usage frequency data.
[0187] Specific operation: Analyzes conversation logs and calculates rewards based on an algorithm.
[0188] Output: Calculated reward data.
[0189] Step 14:
[0190] The server distributes the calculated reward based on the provider's registered account information.
[0191] Input: Remuneration data and provider account information.
[0192] Specific operation: The reward is transferred to the provider's account.
[0193] Output: Rewards deposited into the provider's account.
[0194] The above are the specific processing steps of the program in this system.
[0195] (Application example 1)
[0196] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0197] The present invention relates to a technology and system that provides dialogue simulations to meet the needs of users who want to have real-time dialogues with specific conversation partners. It also aims to solve the problem of simultaneously providing a reward distribution mechanism for dialogue data providers. This allows users to enhance their learning by engaging in dialogues that include specialized knowledge, and data providers to gain new revenue sources.
[0198] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0199] In this invention, the server includes means for collecting conversation data, means for cleaning the collected conversation data and storing it in a database, means for generating a large-scale language model for training using the stored conversation data, means for providing an appropriate AI model based on a specific conversation partner selected by an authenticated user, means for conducting a conversation between the user and the AI model in real time and for the AI model to generate a response based on the user's input, means for calculating and distributing rewards to data providers based on the used conversation data, means for providing an interface for the user to start a dialogue simulation, and means for tracking the user's conversation data and transferring rewards to dialogue data providers through the system. This allows users to simulate a dialogue with a specific conversation partner in real time and effectively learn and exchange information, while also allowing dialogue data providers to receive fair rewards.
[0200] "Conversation Data" means information in the form of audio, text, or video exchanged between Users or between Users and the System.
[0201] "Cleaning" refers to the process of removing noise, inappropriate content, and personal information from collected conversation data.
[0202] "Database" refers to a system that accumulates and manages cleaned conversation data.
[0203] A "large-scale language model" refers to a natural language processing model trained on a massive amount of conversational data.
[0204] "AI model" refers to a large-scale language model tuned based on a specific conversation partner.
[0205] "Real-time" refers to a situation where there is almost no delay and processing or response is immediate.
[0206] "Conversation Partner" refers to a specific expert, celebrity, or AI model thereof with whom the user wishes to have a conversation.
[0207] "Interface" refers to the screen and input means that users use to operate the system.
[0208] "Tracking" refers to the act of following and recording user behavior and conversation data.
[0209] "Reward" refers to money or other consideration paid by the system to the dialogue data provider.
[0210] The present invention is a system that allows users to have real-time conversations with specific conversation partners. The system provides a series of functions including collection, cleaning, and storage of conversation data, generation of AI models, user authentication, real-time conversation execution, and reward distribution.
[0211] System program generation
[0212] Data collection and cleaning
[0213] The server provides an interface for conversation data providers (such as experts and celebrities) to upload audio, text, and video data after they log in. This data is temporarily stored and analyzed to filter out noise, unnecessary parts, and personal information. Once preprocessing is complete, the data is stored in a database as cleaned data.
[0214] Generating AI models
[0215] The server uses the cleaned data to train large-scale language models, which are then tuned to specific conversational partners and set up API endpoints to make them accessible.
[0216] User authentication and conversation simulation
[0217] The terminal (smartphone) provides an authentication interface for users to access and log in to the system. Once authentication is complete, the user is presented with an interface to select a specific conversation partner and begin the dialogue simulation. The user's input is sent to the server, and the corresponding AI model generates a response and sends it back to the terminal. The server performs this process in real time, providing an environment in which the user can continue the dialogue.
[0218] Reward distribution
[0219] The server tracks the user's conversation time and frequency of use, calculates the reward for the conversation data provider, and then processes the payment based on the provider's account information.
[0220] Hardware and software used
[0221] Hardware: A server with a high-performance CPU and GPU, and a smartphone (iOS or Android device)
[0222] Software: Applications running on edge devices (Swift or Kotlin), Python scripts for data analysis and filtering, database management system (MySQL), large-scale language model (OpenAI GPT), API server (Flask or Django)
[0223] Specific examples
[0224] For example, if a user logs in to the system and selects a simulated conversation with a specific expert, the following process will occur: When the user types, "Tell me about your new business idea," the server receives this input and passes it to an AI model trained as a specific expert. The AI model generates a response: "The key to a new business idea is market research and finding a niche." The generated response is sent back to the terminal, and the user can use it to ask further questions.
[0225] Prompt Sentence Examples
[0226] "If a user asks, 'Tell me about a new business idea,' generate a response from an AI trained as a business expert."
[0227] This allows users to enhance their learning by engaging in conversations that include specialized knowledge, and provides conversation data providers with a new source of revenue.
[0228] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0229] Program processing steps
[0230] Step 1:
[0231] Conversation data collection
[0232] Server: When a conversation data provider logs in, it provides an interface for uploading voice data, text data, and video data. Once the provider uploads the data, it is temporarily stored.
[0233] Input: Conversation data (audio, text, video)
[0234] Output: Temporarily saved conversation data
[0235] Specific operation: After authenticating the provider, the server stores the uploaded data in a specific directory on the server.
[0236] Step 2:
[0237] Data Preprocessing
[0238] Server: Analyzes the collected conversation data, filters out noise, unnecessary parts, and personal information, and stores the cleaned data in a database.
[0239] Input: Temporarily saved conversation data
[0240] Output: Cleaned conversation data
[0241] What it does: It uses a Python script to denoise the audio data, sanitize the text, and filter out personal information, then stores it in a MySQL database.
[0242] Step 3:
[0243] Generating AI models
[0244] Server: Trains large-scale language models using cleaned data, tunes AI models based on specific conversational partners, and sets up API endpoints for access.
[0245] Input: Cleaned conversation data
[0246] Output: A tuned AI model
[0247] What it does: Train and tune large-scale language models using Python and TensorFlow. Set up API endpoints using Flask or Django.
[0248] Step 4:
[0249] User authentication and selection
[0250] Terminal (smartphone): Provides an authentication interface for users to access the system through the app and log in. Once authenticated, users can select a specific conversation partner.
[0251] Input: User login information, conversation partner selection
[0252] Output: Preparation for starting the conversation simulation
[0253] Specific operation: The login screen is displayed on the device using Swift or Kotlin, and the user's input information is sent to the server for authentication. After authentication, a list of conversation partners is displayed.
[0254] Step 5:
[0255] Conversation Simulation
[0256] Terminal (smartphone): Displays an interface that allows the user to start a conversation with a selected conversation partner and sends the user's input to the server.
[0257] Server: Passes user input to the AI model, generates a response, and sends it back to the device, which displays the response to the user.
[0258] Input: User's spoken input
[0259] Output: Response by the AI model
[0260] Specific operation: Receives user input and sends it to the server, which passes the input to the AI model and returns the generated response to the mobile device, which displays the response.
[0261] Step 6:
[0262] Reward distribution
[0263] Server: Tracks user conversation time and frequency of use, calculates rewards for conversation data providers, and processes payments.
[0264] Input: Conversation transcript and usage data
[0265] Output: Reward calculation and payment
[0266] Specific operation: A Python script aggregates conversation time and frequency of use, calculates the reward for the provider, and processes the payment to the provider's registered account.
[0267] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0268] The present invention provides a system that collects, cleans, and stores conversation data in a database. It also generates a large-scale language model for training using the stored conversation data, and provides an appropriate AI model based on a specific conversation partner after a user is authenticated. It also conducts real-time conversations between the user and the AI model, and the AI model generates responses based on the user's input. It also includes an emotion engine that recognizes the user's emotions and generates responses adapted to those emotions. This emotion engine can recognize emotions from the user's input voice data and facial expression data. It also includes a means for calculating and distributing rewards to data providers based on the conversation data used.
[0269] Program processing
[0270] 1. Collecting conversation data
[0271] Server: Conversation data providers log in to the system and upload data in the form of audio, text, or video through the conversation data upload interface. This data is temporarily stored on the server.
[0272] 2. Data Preprocessing
[0273] Server: The uploaded conversation data is cleaned. This process includes removing noise and filtering inappropriate content and personal information. The cleaned data is then stored in a database.
[0274] 3. AI model generation
[0275] Server: Trains large-scale language models using cleaned conversation data, trains models specific to specific conversation partners, and sets up API endpoints for access.
[0276] 4. Emotion engine integration
[0277] Server: Integrates an emotion engine into the AI model. The emotion engine contains algorithms for recognizing emotions from user input (text, voice, facial expressions, etc.) and adjusting responses based on the recognized emotions.
[0278] 5. User Authentication and Selection
[0279] Terminal: The user launches the application and accesses the login screen. The user enters their credentials and logs in.
[0280] 6. Conversation Simulation
[0281] Terminal: The user selects a conversation partner through the system's interface and starts a dialogue. The terminal transmits the text and voice data entered by the user to the server.
[0282] Server: Receives user input and passes it to the appropriate AI model and emotion engine. The AI model generates a basic response, and the emotion engine tailors the response based on the user's emotion. For example, if a user makes an input that sounds like they're in distress, the emotion engine will recognize that and generate a response like, "You seem distressed. Is there anything I can help you with?"
[0283] Terminal: displays the adjusted response to the user and allows the user to continue typing.
[0284] 7. Reward Distribution
[0285] Server: Tracks user conversation session data and calculates rewards for data providers based on the time and frequency of use, and pays providers.
[0286] Specific examples
[0287] For example, consider the case where psychological counselor C provides audio data from a session. Counselor C's data is collected, cleaned, and stored in a database. Next, a large-scale language model is trained based on psychological counselor C's data to build an AI model that generates specific responses.
[0288] A user logs into the system and selects "Simulated conversation with psychological counselor C." When the user enters "I've been feeling very stressed lately," the message is sent to the server. The emotion engine analyzes this input and recognizes the user's emotions as "stress" or "anxiety." The AI model generates a basic response of "That must be tough," which the emotion engine refines to "What situations are causing you stress? Tell me about them." This response is ultimately returned to the user.
[0289] This system allows users to have a realistic conversational experience, and allows them to receive psychological support while learning a language. Conversation data providers are also rewarded based on the use of their data.
[0290] The processing flow will be explained below.
[0291] Understood. Below is the process flow broken down into specific steps.
[0292] Step 1:
[0293] Server: Provides an authentication interface for conversation data providers to log in. The conversation data provider logs in by entering the correct authentication information.
[0294] Step 2:
[0295] Server: After successful authentication, the server displays an interface for uploading conversation data to the data provider, allowing the provider to select and upload data in audio, text, or video format.
[0296] Step 3:
[0297] Server: Temporarily stores the uploaded data, checks the format and content of the data, and if the format is correct, sends the data through the cleaning process.
[0298] Step 4:
[0299] Server: The cleaning process removes noise and unnecessary parts from the conversation data and filters out personal information, such as the patient's name and contact details.
[0300] Step 5:
[0301] Server: Organizes the cleaned data and stores it in a database. When storing the data, it categorizes it by conversation topic, person, date and time, etc.
[0302] Step 6:
[0303] Server: Trains a large-scale language model (LLM) using the cleaned conversation data. For example, it trains the model based on the data of psychological counselor C and builds a response generation model specialized for that person.
[0304] Step 7:
[0305] Server: Evaluates the trained model on the test data, adjusts the model parameters if necessary, saves the completed model, and sets up the appropriate API endpoints.
[0306] Step 8:
[0307] Server: Integrates the emotion engine into the AI model. The emotion engine contains algorithms to recognize emotions from user input (text, voice, facial expressions, etc.) and make appropriate adjustments.
[0308] Step 9:
[0309] Terminal: The user launches the application and accesses the login screen. The user enters their credentials and logs in.
[0310] Step 10:
[0311] Server: Checks the user's authentication information, and if authentication is successful, displays the user dashboard, where the user can select the conversation partner they want to simulate (e.g., psychological counselor C).
[0312] Step 11:
[0313] Terminal: Provides an interface for the user to select a simulated conversation partner. The user makes the selection and sends the selection information to the server.
[0314] Step 12:
[0315] Server: To start a simulation based on the user's selection, prepare the corresponding AI model (psychological counselor C's model) and send information to the user's device to optimize performance.
[0316] Step 13:
[0317] Device: Sends text or voice data entered by the user to the server. For example, you might enter, "I've been feeling really stressed lately."
[0318] Step 14:
[0319] Server: Receives user input and passes it to the appropriate AI model. The emotion engine also analyzes this input. The AI model generates a basic response such as "That must be tough," while the emotion engine recognizes emotions such as "stress" and "anxiety" and adjusts the response to "What situations are stressing you out? Tell me about them."
[0320] Step 15:
[0321] Server: Sends the generated response to the user's device.
[0322] Step 16:
[0323] Terminal: Receives the response from the server and displays it in the user interface. The user can then input again if they wish to continue the interaction.
[0324] Step 17:
[0325] Server: After the user's simulation is completed, the server calculates the reward for the data provider based on the simulation time and number of times used, and pays the provider.
[0326] The above is a specific processing flow of the embodiment of this system.
[0327] Example 2
[0328] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0329] Conventional dialogue systems have the following problems. First, the quality of the collected conversation data is inconsistent, resulting in low reliability of the data used for training. Second, they lack the ability to understand the user's emotions and generate responses accordingly, limiting the user experience. Furthermore, the distribution of rewards to data providers is unclear, resulting in a lack of incentive to provide data. These problems must be resolved.
[0330] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0331] In this invention, the server includes means for collecting conversation data, means for cleaning the collected conversation data and storing it in a database, means for generating a large-scale language model trained using the stored conversation data, means for integrating an emotion engine that recognizes user emotions and adjusts responses, and means for calculating and distributing rewards to data providers based on the conversation data used. This makes it possible to generate a reliable large-scale language model using high-quality data and provide appropriate responses according to the user's emotions. Furthermore, transparency of rewards for conversation data providers can be ensured, improving motivation for data provision.
[0332] "Conversational Data" refers to the content of conversations between humans recorded in audio, text, or video format.
[0333] "Means of collection" refers to the mechanism by which user-submitted conversation data is incorporated into the system.
[0334] "Cleaning methods" refers to the process used to remove noise, inappropriate content, and personal information from collected conversation data.
[0335] "Means for storing in a database" refers to the technology used to store the cleaned conversation data as part of a database.
[0336] "Means for generating large-scale language models" refers to a method for using cleaned conversational data to create trained language models based on natural language processing.
[0337] "Means of integrating an emotion engine" refers to the process of incorporating a system into an AI model to recognize emotions from user input data and adjust the model's response based on those emotions.
[0338] "Means for calculating and distributing rewards" refers to a system that calculates and appropriately distributes rewards to data providers based on the amount and frequency of conversation data used.
[0339] "User authentication" refers to the process of verifying a user's identity when logging into a system.
[0340] "Conversation Partner" refers to a virtual or real entity that a user selects within the system with whom to interact.
[0341] "Real-time" means that user input is processed almost instantly and responses are returned immediately.
[0342] The present invention is a system that collects, cleans, and stores conversation data in a database. Furthermore, the system generates a large-scale language model for training using the stored conversation data, and provides an appropriate AI model based on a specific conversation partner after the user is authenticated. The system conducts real-time conversations between the user and the AI model, and the AI model generates responses based on the user's input. The system also includes an emotion engine that recognizes the user's emotions and generates responses adapted to those emotions. The emotion engine can recognize emotions from the user's input voice data and facial expression data. The system also includes a means for calculating and distributing rewards to data providers based on the conversation data used.
[0343] Hardware and software used
[0344] The system uses the following hardware and software:
[0345] Server: A server with high performance data storage and computing power (e.g., AWS EC2 instance)
[0346] Database: A relational database (e.g., Amazon RDS, MySQL) to store conversation data.
[0347] Deep learning frameworks for training language models (e.g., TensorFlow, PyTorch)
[0348] Emotion recognition engine: Software for analyzing voice and facial expression data (e.g., OpenCV, emotionAPI)
[0349] User interface: Web or mobile application for data upload and conversation simulation
[0350] Collecting and cleaning conversation data
[0351] The server provides an interface for conversation data providers to log in and upload conversation data in the form of audio, text, or video. Providers upload the data, which is temporarily stored on the server. The server then cleans the uploaded conversation data, removing noise and filtering inappropriate content and personal information, and stores the cleaned data in a database.
[0352] Generating large-scale language models
[0353] The server uses the cleaned conversation data to train large-scale language models, generates models tailored to specific conversation partners, and sets up API endpoints to make them accessible. Training is done using deep learning frameworks such as TensorFlow and PyTorch.
[0354] Emotion engine integration
[0355] The server integrates an emotion engine into the AI model. The emotion engine recognizes emotions from user input data (text, voice, facial expressions) and adjusts responses based on the results. The emotion recognition algorithm evaluates the user's emotional state in real time and generates appropriate responses.
[0356] User authentication and conversation initiation
[0357] The user launches the application on their device and enters their authentication information on a login screen. If authentication is successful, an interface appears where they can select a conversation partner and begin the dialogue. The user inputs text or voice, and the data is sent to the server. The server passes the received input to an AI model and emotion engine, generating a response in real time. The response is displayed on the device, and the user can continue the conversation.
[0358] Reward distribution
[0359] The server tracks users' conversation data and calculates rewards for data providers based on it. Rewards are calculated based on usage time and frequency and distributed appropriately. This increases transparency for data providers and strengthens their motivation to provide data.
[0360] Specific examples
[0361] For example, consider a case where a user simulates a conversation with a psychological counselor. The user inputs, "I've been feeling very stressed lately," and the message is sent to the server. The emotion engine analyzes this input and recognizes the user's emotion as "stress." The AI model generates a basic response, "That must be tough," which the emotion engine then refines by saying, "What situations are causing you stress? Tell me about them." This response is ultimately returned to the user.
[0362] Prompt Sentence Examples
[0363] Example prompt: "Tell me about the latest robotics technology."
[0364] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0365] Step 1:
[0366] Conversation data collection
[0367] Input: Audio, text, or video files provided by the conversation data provider
[0368] Server: A conversation data provider logs in to the system and uploads conversation data through the data upload interface. When the provider selects the data and clicks the upload button, the data is sent to the server and temporarily stored.
[0369] Output: Temporarily saved conversation data
[0370] Specific behavior:
[0371] The user clicks the "Upload conversation data" button.
[0372] The selected file is sent to the server and saved in a temporary folder.
[0373] Step 2:
[0374] Data Preprocessing
[0375] Input: Temporarily saved conversation data
[0376] Server: The uploaded conversation data is cleaned. This process includes removing noise and filtering inappropriate content and personal information. The cleaned data is then stored in a database.
[0377] Output: Cleaned conversation data
[0378] Specific behavior:
[0379] The server adds the uploaded data to a processing queue.
[0380] A cleaning algorithm steps through the data in the queue and applies a noise reduction filter.
[0381] Apply a set of personal information filtering rules to automatically remove inappropriate content.
[0382] Store clean data in the database.
[0383] Step 3:
[0384] AI model generation
[0385] Input: Cleaned conversation data
[0386] Server: Trains large-scale language models using cleaned conversation data, generates models tailored to specific conversation partners, and sets up API endpoints to make them accessible.
[0387] Output: A trained large-scale language model
[0388] Specific behavior:
[0389] The server retrieves the cleaned data from the database.
[0390] Train the model using a deep learning framework (e.g., PyTorch, TensorFlow).
[0391] Save the parameters of the model once it has been trained.
[0392] Create an API endpoint and deploy the model.
[0393] Step 4:
[0394] Emotion engine integration
[0395] Input: A trained large-scale language model
[0396] Server: Integrates an emotion engine into the AI model. The emotion engine contains algorithms for recognizing emotions from user input data (text, voice, facial expressions) and adjusting responses based on the results.
[0397] Output: AI model with integrated emotion engine
[0398] Specific behavior:
[0399] Emotion recognition algorithms perform text and voice analysis.
[0400] Assessing the user's emotional state (e.g., happy, sad, angry) in real time.
[0401] The basic responses generated by the AI model are fine-tuned based on the emotion recognition results to generate the optimal response.
[0402] Step 5:
[0403] User authentication and selection
[0404] Input: Authentication information (email address, password)
[0405] On the device: The user launches the application and accesses the login screen. The user enters their email address and password for authentication. If successful, the conversation partner selection screen is displayed.
[0406] Output: List of conversation partners after successful authentication
[0407] Specific behavior:
[0408] The user launches the application and enters their login ID and password.
[0409] The server verifies the authentication information and, if it matches, starts the session.
[0410] After successful login, a list of conversation partners will be displayed on your device.
[0411] Step 6:
[0412] Conversation Simulation
[0413] Input: User text or voice input
[0414] Terminal: The user selects a conversation partner through the system's interface and starts a conversation. When the user inputs text or voice, the data is sent to the server.
[0415] Server: Receives user input and passes it to the appropriate AI model and emotion engine. The AI model generates a basic response, and the emotion engine adjusts the response based on the user's emotion.
[0416] Output: Adjusted AI response
[0417] Specific behavior:
[0418] The user selects "Psychological Counselor C" and clicks the "Start Dialogue" button.
[0419] The user types "I've been feeling really stressed lately" into the input field and submits.
[0420] The server receives this input and sends it to the AI model.
[0421] The model generates the response "That's tough," which the emotion engine refines to "What situations are stressing you out? Tell me about them."
[0422] The response is returned to the terminal and displayed to the user.
[0423] Step 7:
[0424] Reward distribution
[0425] Input: conversation session data (usage time, usage frequency)
[0426] Server: Tracks user conversation records and calculates rewards for data providers based on them. Calculates rewards based on the time and frequency of use and pays providers.
[0427] Output: Calculated reward points and payment processing
[0428] Specific behavior:
[0429] The server records data for each conversation session.
[0430] A reward calculation algorithm calculates points based on the duration and frequency of use per session.
[0431] Rewards are transferred to conversation data providers through a payment processing service.
[0432] (Application example 2)
[0433] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0434] Customer service in traditional brick-and-mortar stores is heavily dependent on the experience and skills of store staff, making it difficult to maintain consistent service quality. Furthermore, in situations where flexible responses based on customer emotions are required, it is difficult for humans alone to recognize and respond 100% accurately. Furthermore, the appropriate management and effective use of collected conversation data is an issue.
[0435] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting conversation data, means for cleaning the collected conversation data and storing it in a database, means for generating a large-scale language model to be trained using the stored conversation data, means for providing an appropriate AI model based on a specific conversation partner selected by an authenticated user, means for conducting a real-time conversation between the user and the AI model and for the AI model to generate a response based on the user's input, means for generating a response based on the user's emotions using an emotion recognition engine, means for analyzing collected audio and video data using smart glasses and providing appropriate responses to store clerks in real time, and means for calculating and distributing rewards to data providers based on the used conversation data. This enables real-time responses that take customer emotions into consideration and uniform service quality.
[0436] "Conversation Data" means records of audio, text, and video communications between users or between users and the system.
[0437] "Cleaning" is the process of removing noise and organizing data while removing inappropriate content and personal information.
[0438] A "database" is a collection of information that stores collected data in an organized manner and makes it easily searchable and accessible.
[0439] A "large-scale language model" is an artificial intelligence model trained using vast amounts of conversational data to generate natural-sounding language responses based on user input.
[0440] An "AI model" is an artificial intelligence algorithm and its implementation designed to address a specific task or application.
[0441] An "emotion recognition engine" analyzes emotions from a user's voice and video data and adjusts responses based on those emotions.
[0442] "Smart glasses" are a type of wearable device used to acquire and display visual and audio information, and have the ability to present analysis results to users in real time.
[0443] "Reward" means compensation distributed to data providers based on the use of the conversation data they provide.
[0444] "Data providers" are individuals or organizations that upload conversation data to the system.
[0445] The present invention is a system that aims to improve the quality and efficiency of customer support in brick-and-mortar stores, and includes means for collecting and cleaning conversation data, generating large-scale language models, recognizing emotions, and distributing rewards. The following describes in detail the embodiments of the present invention.
[0446] System Program
[0447] 1. Collecting conversation data
[0448] The server collects conversation data through smart glasses worn by store staff. The smart glasses are equipped with a microphone to collect voice data and a camera to capture facial expression data. The collected data is both audio and video.
[0449] 2. Data Preprocessing
[0450] The server receives the raw conversation data sent by the smart glasses and cleans it. The cleaning process includes removing noise and filtering inappropriate content and personal information, improving the quality of the data before storing it in a database.
[0451] 3. Generating large-scale language models
[0452] The server uses the cleaned conversation data to train a large-scale language model, which is trained specifically for specific conversational partners (in this case, the store clerk and the customer) and made accessible through an API, using Python and AI frameworks such as TensorFlow.
[0453] 4. Emotion Recognition Integration
[0454] The server integrates a large-scale language model with an emotion recognition engine, which uses OpenCV and TensorFlow to analyze user emotions from audio and video data. The analysis results are reflected in the generated response.
[0455] 5. User Authentication and Response Generation
[0456] The terminal (smart glasses) performs personal authentication when the store clerk logs in. After logging in, when the store clerk interacts with the customer, voice input and video data are sent to the server. The server receives this data and passes it to an AI model and emotion recognition engine to generate an appropriate response.
[0457] 6. Real-time response
[0458] The server sends the generated response to the smart glasses in real time and displays it to the store clerk. This allows the store clerk to immediately provide an appropriate response to the customer. For example, if a customer asks, "I've been feeling stressed lately," the smart glasses will respond, "That's tough. What situations are making you feel stressed?"
[0459] 7. Reward Distribution
[0460] The server calculates and distributes rewards to data providers (store clerks and stores) based on the collected conversation data. Rewards are determined based on the frequency and quality of data usage.
[0461] Hardware and software used
[0462] Hardware: Smart glasses (e.g., Google Glass Enterprise Edition 2)
[0463] Server software: AWS Lambda, S3, EC2
[0464] Preprocessing and training: Python scripts, TensorFlow
[0465] Emotion recognition: OpenCV, TensorFlow
[0466] Authentication system: OAuth 2.0, Firebase Authentication
[0467] Examples of prompt statements
[0468] For example, if a customer asks, "I've been feeling stressed lately, especially at work...what should I do?" the system might respond with:
[0469] "That's tough. What situations are stressing you out?"
[0470] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0471] Step 1:
[0472] A user (a store clerk) wears smart glasses and serves customers in a store. The smart glasses collect audio and video data in real time and send it to a server. The input is conversation data (audio and video) between the customer and the store clerk, which is then sent to the server.
[0473] Step 2:
[0474] The server receives the transmitted conversation data and performs cleaning processes such as noise removal, filtering of inappropriate content and personal information. The data is then cleaned and stored in a database. The input is raw conversation data, and the output is cleaned data.
[0475] Step 3:
[0476] The server uses the cleaned conversation data to train a large-scale language model. It uses an AI framework such as TensorFlow to generate a model specialized for a specific conversation partner. The input is the cleaned conversation data, and the output is the trained large-scale language model.
[0477] Step 4:
[0478] The server integrates an emotion recognition engine with the trained large-scale language model. This engine uses OpenCV and TensorFlow to recognize user emotions from audio and video data. The input is audio and video data, and the output is recognized emotional information and response adjustments based on it.
[0479] Step 5:
[0480] The smart glasses at the terminal authenticate the store clerk by logging in to the system. During this process, the user's authentication information is sent to the server and verified. The input is the store clerk's authentication information, and the output is the authentication success or failure status.
[0481] Step 6:
[0482] The server receives real-time audio and video data when an authenticated store associate interacts with a customer and passes it to an AI model and emotion recognition engine. The AI model and emotion recognition engine then generate an appropriate response based on the user's input data. The input is real-time conversation data with the customer, and the output is the generated appropriate response.
[0483] Step 7:
[0484] The server sends the generated response to the smart glasses in real time and presents it to the store clerk. The store clerk responds to the customer based on the response. The input is the generated response, and the output is the response displayed on the smart glasses.
[0485] Step 8:
[0486] The server tracks all conversation sessions and calculates and distributes rewards to data providers based on the conversation data used. The input is the conversation session data, and the output is the reward calculation and distribution results.
[0487] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0488] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0489] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0490] [Second embodiment]
[0491] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0492] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0493] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0494] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0495] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0496] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0497] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0498] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0499] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0500] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0501] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0502] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0503] The present invention provides a system that collects, cleans, and stores conversation data in a database, generates a large-scale language model for training using the stored conversation data, and provides an appropriate AI model based on a specific conversation partner after a user is authenticated. The system also conducts real-time conversations between the user and the AI model, and the AI model generates responses based on the user's input. The system also includes a means for calculating and distributing rewards to data providers based on the conversation data used.
[0504] Program processing
[0505] 1. Collecting conversation data
[0506] Server: Provides an interface for conversation data providers (celebrities and professionals) to upload their conversation data after logging in. When providers upload audio, text, or video data, the data is temporarily stored.
[0507] 2. Data Preprocessing
[0508] Server: Analyzes the received conversation data, removes noise and unnecessary parts, and filters out personal information. The preprocessed data is saved in the database as cleaned data.
[0509] 3. AI model generation
[0510] Server: Uses the cleaned data to train large-scale language models, which are then tuned to specific conversational partners and set up API endpoints to make them accessible.
[0511] 4. User Authentication and Selection
[0512] Terminal: Provides an authentication interface for users to access the system through the application and log in. Once authenticated, users can select a specific conversation partner from a list.
[0513] 5. Conversation Simulation
[0514] Terminal: The terminal displays an input form on the interface to start a conversation with a conversation partner of the user's choice. The user enters the information and sends it to the server.
[0515] Server: Receives user input and passes it to the appropriate AI model, which generates a response and sends it back to the device.
[0516] Terminal: Displays responses from the server and provides an environment in which the user can continue the dialogue.
[0517] 6. Reward Distribution
[0518] Server: Tracks the user's conversation time and frequency of use, and calculates the reward for the conversation data provider. Once the reward is calculated, payment is processed based on the provider's account information.
[0519] Specific examples
[0520] For example, consider the case where a famous actor B provides audio data about his or her stage experience. This data is collected, cleaned to remove personal information, and then stored in a database. Next, a large-scale language model is trained based on Actor B's audio data to generate an AI model that can converse with Actor B.
[0521] A user logs into this system and selects "Simulated dialogue with actor B." When the user inputs "What was your most recent stage challenge?" to actor B, the input is sent to the server, and the AI model generates a response such as "My most recent stage challenge was trying out a new role," and sends it back to the user.
[0522] When the conversation ends, the server calculates the reward for Actor B based on the time and frequency of use and transfers it to the system. In this way, users can advance their language learning through a realistic dialogue experience, and the provider of the conversation data can also earn rewards.
[0523] This system not only allows language learners to stay motivated and improve their skills effectively, but also allows conversation data providers to earn revenue in new ways.
[0524] The processing flow will be explained below.
[0525] Understood. Below is the process flow broken down into specific steps.
[0526] Step 1:
[0527] Server: Provides an authentication interface for conversation data providers (e.g., famous actor B) to log in. The conversation data provider logs in by entering the correct authentication information.
[0528] Step 2:
[0529] Server: After successful authentication, the server displays a conversation data upload interface to the conversation data provider, where the provider can select and upload conversation data in audio, text, or video format.
[0530] Step 3:
[0531] Server: Temporarily stores the uploaded data, checks the format and content of the data, and if the format is correct, sends the data to the cleaning process.
[0532] Step 4:
[0533] Server: The cleaning process removes noise and unnecessary parts from the conversation data and filters out personal information, such as the patient's name and contact details.
[0534] Step 5:
[0535] Server: Organizes the cleaned data and stores it in a database. When storing the data, it categorizes it by conversation topic, person, date and time, etc.
[0536] Step 6:
[0537] Server: Uses the cleaned conversation data to train a large-scale language model (LLM). For example, it trains the model based on actor B's data and builds a response generation model specialized for that person.
[0538] Step 7:
[0539] Server: Evaluates the trained model on the test data, adjusts the model parameters if necessary, saves the completed model, and sets up the appropriate API endpoints.
[0540] Step 8:
[0541] Terminal: The user launches the application and accesses the login screen. The user enters their credentials and logs in.
[0542] Step 9:
[0543] Server: Checks the user's authentication information, and if authentication is successful, displays the user dashboard, where the user can select the conversation partner they want to simulate (e.g., Actor B).
[0544] Step 10:
[0545] Terminal: Displays an interface for the user to select a simulated conversation partner. The user makes the selection and sends the selection information to the server.
[0546] Step 11:
[0547] Server: Based on the user's selection, prepares the corresponding AI model (e.g., actor B's model) and sends interface information to the terminal to start the conversation.
[0548] Step 12:
[0549] Terminal: Sends a text message typed by the user to the server. For example, "What's your latest stage challenge?"
[0550] Step 13:
[0551] Server: Receives user input and passes it to the appropriate AI model, which then generates a response, for example, "My most recent stage challenge was taking on a new role."
[0552] Step 14:
[0553] Server: Sends the generated response to the user's device.
[0554] Step 15:
[0555] Terminal: Receives the response from the server and displays it in the user interface. If the user wants to continue the interaction, they can input again.
[0556] Step 16:
[0557] Server: After the user's simulation is completed, the server calculates the reward for the conversation data provider based on the simulation time and number of times used.
[0558] Step 17:
[0559] Server: Carries out the payment procedures to distribute the calculated reward to the conversation data provider. The reward is transferred based on the provider's account information.
[0560] The above is the specific processing flow of this system.
[0561] Example 1
[0562] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0563] Conventional conversation data collection and processing systems have insufficient quality control of collected data, particularly due to the inclusion of noise and personal information, which limits its use. Furthermore, when training AI models using conversation data, it is difficult to adjust them to specific conversation partners, making it difficult to provide real-time responses to users. Furthermore, reward distribution to data providers is sometimes not transparent and efficient, resulting in a lack of incentives for providers.
[0564] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0565] In this invention, the server includes a means for collecting conversation data, a means for cleaning the collected conversation data and storing it in a database, and a means for generating a model to be trained using the stored conversation data. This enables appropriate preprocessing of the data and the generation of an AI model tailored to a specific conversation partner using high-quality training data. The server also includes a means for selecting a specific conversation partner after the user is authenticated, a means for conducting a conversation between the user and the model and for the model to generate a response based on the user's input, and a means for calculating and distributing rewards to data providers based on the used conversation data. This allows users to receive appropriate responses in real time and distributes rewards fairly and quickly to data providers, improving the efficiency and satisfaction of the entire system.
[0566] "Conversation data" refers to audio data, text data, or video data provided to the system by users or conversation data providers.
[0567] "Cleaning" is the process of removing noise, inappropriate content, personal information, etc. from collected conversation data to improve data quality before storing it in a database.
[0568] A "database" is a storage within the system that efficiently and securely stores cleaned conversation data for later use as training data.
[0569] "Training" is the process of using cleaned conversational data to train an AI model and improve its pattern recognition capabilities and prediction accuracy.
[0570] A "model" is a collection of AI structures and algorithms trained on conversational data that are used to generate appropriate responses to user input.
[0571] A "specific conversation partner" is a virtual conversation partner for a dialogue that is selected by the user based on various categories and profiles provided on the system.
[0572] "Real-time" means that there is a very short delay between when a user makes an input and when the AI model returns a response, and refers to a situation in which the interaction takes place immediately.
[0573] A "response" is a response, such as text or voice, that an AI model generates based on user input.
[0574] "Remuneration" is the compensation calculated and paid to the conversation data provider based on the degree and frequency of use of the provided data.
[0575] "Distribution" is the process of appropriately paying the calculated reward based on the registered information of the conversation data provider.
[0576] "Interface" refers to the screens, input forms, and operating means used by conversation data providers when uploading data or by users when accessing and operating the system.
[0577] To implement this invention, the server, terminal, and user each play specific roles. The main hardware used includes a database server, web server, and user terminal (PC or smartphone), while the software required includes a framework for training large-scale language models (e.g., TensorFlow, PyTorch), database software (e.g., MySQL, PostgreSQL), and a web service API.
[0578] The server provides a means for collecting conversation data. Specifically, it provides a user interface for conversation data providers to log in and upload audio data, text data, and video data. The device supports a means for providers to upload data, allowing them to select and send data in various formats (e.g., .mp3, .txt, .mp4). Uploaded data is temporarily stored on the server.
[0579] The server analyzes the collected conversation data, removes noise and unnecessary parts, and filters out personal information to clean it. The cleaned data is then stored in a database, which later serves as the foundation for training AI models.
[0580] The server then uses the cleaned conversation data to train a large-scale language model using a large-scale language model framework, processing the data in batches to learn the model weights, resulting in an AI model that corresponds to a specific conversation partner. The model is then made accessible through an API endpoint.
[0581] To access the system, a user uses a terminal to enter a username and password into a login interface. After successful authentication, the server provides the user with a list of specific available conversation partners from which the user can select a conversation partner.
[0582] The device then displays an interface for the user to converse with the selected conversation partner. The user enters a question or message into the input form and clicks the send button. For example, a question such as "What was your most recent stage challenge?" is sent to the server.
[0583] The server receives user input, passes it to the trained AI model, and generates a response, which is then sent back to the device, where it is displayed to the user, allowing the user to continuously engage in real-time dialogue with the AI model.
[0584] Finally, the server tracks the duration and frequency of users' conversation data usage and calculates rewards for conversation data providers based on that data. The calculated rewards are distributed based on the providers' registered account information.
[0585] As a concrete example, consider a case where a user selects a dialogue simulation with actor B and asks, "What was the most moving moment on stage?" This question is transmitted to the AI model via the server, and the AI model generates a response such as, "The most moving moment on stage was when the audience shed tears," and sends it back to the user.
[0586] This system allows users to advance their language learning through realistic conversational experiences, and enables conversation data providers to earn revenue in new ways.
[0587] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0588] Step 1:
[0589] The server verifies the login of the conversation data provider and, after successful authentication, provides an interface for data upload.
[0590] Input: Provider username and password.
[0591] Specific operation: Query the database to verify the provider's authentication information, and if authentication is successful, start a login session.
[0592] Output: Login session token and data upload interface.
[0593] Step 2:
[0594] Using the terminal, the conversation data provider selects audio data, text data, or video data and clicks the upload button.
[0595] Input: Data files selected by the provider (e.g. .mp3, .txt, .mp4).
[0596] Specific operation: A file selection dialog is displayed, and after the provider selects a file, they press the upload button.
[0597] Output: The selected data file is sent to the server.
[0598] Step 3:
[0599] The server stores the uploaded data in temporary storage and checks the integrity of the data.
[0600] Input: Uploaded data file.
[0601] What it does: Saves the file to temporary storage and verifies the data integrity by checking the file format and the first few bytes of the data.
[0602] Output: Data file with integrity checked.
[0603] Step 4:
[0604] The server analyzes the received conversation data, removes noise and unnecessary parts, and filters out personal information to clean it.
[0605] Input: Data file whose integrity has been checked.
[0606] Specific operation: For audio data, a noise reduction algorithm is applied, and for text data, filtering is performed to mask specific keywords.
[0607] Output: Cleaned data.
[0608] Step 5:
[0609] The server organizes and stores the cleaned data in a database.
[0610] Input: Cleaned data.
[0611] What it does: Inserts data and creates indexes according to the appropriate schema in the database.
[0612] Output: Cleaned data stored in a database.
[0613] Step 6:
[0614] The server uses the cleaned conversational data to train a large-scale language model.
[0615] Input: Cleaned data in the database.
[0616] Specific behavior: Use a large-scale language model framework (e.g., TensorFlow, PyTorch) to process data in batches and update the model weights.
[0617] Output: A trained AI model.
[0618] Step 7:
[0619] The server tunes the trained model to correspond to specific conversation partners and sets up an API endpoint to make it accessible externally.
[0620] Input: A trained AI model.
[0621] Specific actions: Adjust model parameters, configure API endpoints, and deploy.
[0622] Output: An accessible API endpoint for the AI model.
[0623] Step 8:
[0624] The device presents the user with a login interface, prompting them to enter their username and password, and upon successful authentication, providing an interface for selecting a specific conversation partner.
[0625] Input: Username and Password.
[0626] Specific behavior: The user enters information into the login form and clicks the login button.
[0627] Output: Conversation partner selection interface for authenticated user.
[0628] Step 9:
[0629] The terminal displays a screen for starting a conversation based on the conversation partner selected by the user.
[0630] Input: User's selected conversation partner information.
[0631] Specific behavior: Displays a conversation interface and provides an input form.
[0632] Output: A form for the user to enter.
[0633] Step 10:
[0634] The user enters a question or message into the input form and clicks the send button.
[0635] Input: User question or message (e.g., "What was your most recent stage challenge?").
[0636] Specific behavior: The user enters text and clicks the submit button.
[0637] Output: The entered question or message is sent to the server.
[0638] Step 11:
[0639] The server receives user input and passes it to a trained AI model to generate a response.
[0640] Input: The user's typed question or message.
[0641] What it does: Tokenizes the user's message and passes it to the AI model for processing.
[0642] Output: The generated response of the AI model.
[0643] Step 12:
[0644] The terminal displays responses from the server to the user, providing an interface that allows the conversation to continue.
[0645] Input: The response sent back from the server.
[0646] Specific behavior: Displays the response text and persists the interface so the user can enter it again.
[0647] Output: The response displayed to the user and an input form for further interaction.
[0648] Step 13:
[0649] The server tracks the time and frequency of users' conversation data usage and calculates rewards for conversation data providers based on that data.
[0650] Input: User conversation time, usage frequency data.
[0651] Specific operation: Analyzes conversation logs and calculates rewards based on an algorithm.
[0652] Output: Calculated reward data.
[0653] Step 14:
[0654] The server distributes the calculated reward based on the provider's registered account information.
[0655] Input: Remuneration data and provider account information.
[0656] Specific operation: The reward is transferred to the provider's account.
[0657] Output: Rewards deposited into the provider's account.
[0658] The above are the specific processing steps of the program in this system.
[0659] (Application example 1)
[0660] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0661] The present invention relates to a technology and system that provides dialogue simulations to meet the needs of users who want to have real-time dialogues with specific conversation partners. It also aims to solve the problem of simultaneously providing a reward distribution mechanism for dialogue data providers. This allows users to enhance their learning by engaging in dialogues that include specialized knowledge, and data providers to gain new revenue sources.
[0662] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0663] In this invention, the server includes means for collecting conversation data, means for cleaning the collected conversation data and storing it in a database, means for generating a large-scale language model for training using the stored conversation data, means for providing an appropriate AI model based on a specific conversation partner selected by an authenticated user, means for conducting a conversation between the user and the AI model in real time and for the AI model to generate a response based on the user's input, means for calculating and distributing rewards to data providers based on the used conversation data, means for providing an interface for the user to start a dialogue simulation, and means for tracking the user's conversation data and transferring rewards to dialogue data providers through the system. This allows users to simulate a dialogue with a specific conversation partner in real time and effectively learn and exchange information, while also allowing dialogue data providers to receive fair rewards.
[0664] "Conversation Data" means information in the form of audio, text, or video exchanged between Users or between Users and the System.
[0665] "Cleaning" refers to the process of removing noise, inappropriate content, and personal information from collected conversation data.
[0666] "Database" refers to a system that accumulates and manages cleaned conversation data.
[0667] A "large-scale language model" refers to a natural language processing model trained on a massive amount of conversational data.
[0668] "AI model" refers to a large-scale language model tuned based on a specific conversation partner.
[0669] "Real-time" refers to a situation where there is almost no delay and processing or response is immediate.
[0670] "Conversation Partner" refers to a specific expert, celebrity, or AI model thereof with whom the user wishes to have a conversation.
[0671] "Interface" refers to the screen and input means that users use to operate the system.
[0672] "Tracking" refers to the act of following and recording user behavior and conversation data.
[0673] "Reward" refers to money or other consideration paid by the system to the dialogue data provider.
[0674] The present invention is a system that allows users to have real-time conversations with specific conversation partners. The system provides a series of functions including collection, cleaning, and storage of conversation data, generation of AI models, user authentication, real-time conversation execution, and reward distribution.
[0675] System program generation
[0676] Data collection and cleaning
[0677] The server provides an interface for conversation data providers (such as experts and celebrities) to upload audio, text, and video data after they log in. This data is temporarily stored and analyzed to filter out noise, unnecessary parts, and personal information. Once preprocessing is complete, the data is stored in a database as cleaned data.
[0678] Generating AI models
[0679] The server uses the cleaned data to train large-scale language models, which are then tuned to specific conversational partners and set up API endpoints to make them accessible.
[0680] User authentication and conversation simulation
[0681] The terminal (smartphone) provides an authentication interface for users to access and log in to the system. Once authentication is complete, the user is presented with an interface to select a specific conversation partner and begin the dialogue simulation. The user's input is sent to the server, and the corresponding AI model generates a response and sends it back to the terminal. The server performs this process in real time, providing an environment in which the user can continue the dialogue.
[0682] Reward distribution
[0683] The server tracks the user's conversation time and frequency of use, calculates the reward for the conversation data provider, and then processes the payment based on the provider's account information.
[0684] Hardware and software used
[0685] Hardware: A server with a high-performance CPU and GPU, and a smartphone (iOS or Android device)
[0686] Software: Applications running on edge devices (Swift or Kotlin), Python scripts for data analysis and filtering, database management system (MySQL), large-scale language model (OpenAI GPT), API server (Flask or Django)
[0687] Specific examples
[0688] For example, if a user logs in to the system and selects a simulated conversation with a specific expert, the following process will occur: When the user types, "Tell me about your new business idea," the server receives this input and passes it to an AI model trained as a specific expert. The AI model generates a response: "The key to a new business idea is market research and finding a niche." The generated response is sent back to the terminal, and the user can use it to ask further questions.
[0689] Prompt Sentence Examples
[0690] "If a user asks, 'Tell me about a new business idea,' generate a response from an AI trained as a business expert."
[0691] This allows users to enhance their learning by engaging in conversations that include specialized knowledge, and provides conversation data providers with a new source of revenue.
[0692] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0693] Program processing steps
[0694] Step 1:
[0695] Conversation data collection
[0696] Server: When a conversation data provider logs in, it provides an interface for uploading voice data, text data, and video data. Once the provider uploads the data, it is temporarily stored.
[0697] Input: Conversation data (audio, text, video)
[0698] Output: Temporarily saved conversation data
[0699] Specific operation: After authenticating the provider, the server stores the uploaded data in a specific directory on the server.
[0700] Step 2:
[0701] Data Preprocessing
[0702] Server: Analyzes the collected conversation data, filters out noise, unnecessary parts, and personal information, and stores the cleaned data in a database.
[0703] Input: Temporarily saved conversation data
[0704] Output: Cleaned conversation data
[0705] What it does: It uses a Python script to denoise the audio data, sanitize the text, and filter out personal information, then stores it in a MySQL database.
[0706] Step 3:
[0707] Generating AI models
[0708] Server: Trains large-scale language models using cleaned data, tunes AI models based on specific conversational partners, and sets up API endpoints for access.
[0709] Input: Cleaned conversation data
[0710] Output: A tuned AI model
[0711] What it does: Train and tune large-scale language models using Python and TensorFlow. Set up API endpoints using Flask or Django.
[0712] Step 4:
[0713] User authentication and selection
[0714] Terminal (smartphone): Provides an authentication interface for users to access the system through the app and log in. Once authenticated, users can select a specific conversation partner.
[0715] Input: User login information, conversation partner selection
[0716] Output: Preparation for starting the conversation simulation
[0717] Specific operation: The login screen is displayed on the device using Swift or Kotlin, and the user's input information is sent to the server for authentication. After authentication, a list of conversation partners is displayed.
[0718] Step 5:
[0719] Conversation Simulation
[0720] Terminal (smartphone): Displays an interface that allows the user to start a conversation with a selected conversation partner and sends the user's input to the server.
[0721] Server: Passes user input to the AI model, generates a response, and sends it back to the device, which displays the response to the user.
[0722] Input: User's spoken input
[0723] Output: Response by the AI model
[0724] Specific operation: Receives user input and sends it to the server, which passes the input to the AI model and returns the generated response to the mobile device, which displays the response.
[0725] Step 6:
[0726] Reward distribution
[0727] Server: Tracks user conversation time and frequency of use, calculates rewards for conversation data providers, and processes payments.
[0728] Input: Conversation transcript and usage data
[0729] Output: Reward calculation and payment
[0730] Specific operation: A Python script aggregates conversation time and frequency of use, calculates the reward for the provider, and processes the payment to the provider's registered account.
[0731] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0732] The present invention provides a system that collects, cleans, and stores conversation data in a database. It also generates a large-scale language model for training using the stored conversation data, and provides an appropriate AI model based on a specific conversation partner after a user is authenticated. It also conducts real-time conversations between the user and the AI model, and the AI model generates responses based on the user's input. It also includes an emotion engine that recognizes the user's emotions and generates responses adapted to those emotions. This emotion engine can recognize emotions from the user's input voice data and facial expression data. It also includes a means for calculating and distributing rewards to data providers based on the conversation data used.
[0733] Program processing
[0734] 1. Collecting conversation data
[0735] Server: Conversation data providers log in to the system and upload data in the form of audio, text, or video through the conversation data upload interface. This data is temporarily stored on the server.
[0736] 2. Data Preprocessing
[0737] Server: The uploaded conversation data is cleaned. This process includes removing noise and filtering inappropriate content and personal information. The cleaned data is then stored in a database.
[0738] 3. AI model generation
[0739] Server: Trains large-scale language models using cleaned conversation data, trains models specific to specific conversation partners, and sets up API endpoints for access.
[0740] 4. Emotion engine integration
[0741] Server: Integrates an emotion engine into the AI model. The emotion engine contains algorithms for recognizing emotions from user input (text, voice, facial expressions, etc.) and adjusting responses based on the recognized emotions.
[0742] 5. User Authentication and Selection
[0743] Terminal: The user launches the application and accesses the login screen. The user enters their credentials and logs in.
[0744] 6. Conversation Simulation
[0745] Terminal: The user selects a conversation partner through the system's interface and starts a dialogue. The terminal transmits the text and voice data entered by the user to the server.
[0746] Server: Receives user input and passes it to the appropriate AI model and emotion engine. The AI model generates a basic response, and the emotion engine tailors the response based on the user's emotion. For example, if a user makes an input that sounds like they're in distress, the emotion engine will recognize that and generate a response like, "You seem distressed. Is there anything I can help you with?"
[0747] Terminal: displays the adjusted response to the user and allows the user to continue typing.
[0748] 7. Reward Distribution
[0749] Server: Tracks user conversation session data and calculates rewards for data providers based on the time and frequency of use, and pays providers.
[0750] Specific examples
[0751] For example, consider the case where psychological counselor C provides audio data from a session. Counselor C's data is collected, cleaned, and stored in a database. Next, a large-scale language model is trained based on psychological counselor C's data to build an AI model that generates specific responses.
[0752] A user logs into the system and selects "Simulated conversation with psychological counselor C." When the user enters "I've been feeling very stressed lately," the message is sent to the server. The emotion engine analyzes this input and recognizes the user's emotions as "stress" or "anxiety." The AI model generates a basic response of "That must be tough," which the emotion engine refines to "What situations are causing you stress? Tell me about them." This response is ultimately returned to the user.
[0753] This system allows users to have a realistic conversational experience, and allows them to receive psychological support while learning a language. Conversation data providers are also rewarded based on the use of their data.
[0754] The processing flow will be explained below.
[0755] Understood. Below is the process flow broken down into specific steps.
[0756] Step 1:
[0757] Server: Provides an authentication interface for conversation data providers to log in. The conversation data provider logs in by entering the correct authentication information.
[0758] Step 2:
[0759] Server: After successful authentication, the server displays an interface for uploading conversation data to the data provider, allowing the provider to select and upload data in audio, text, or video format.
[0760] Step 3:
[0761] Server: Temporarily stores the uploaded data, checks the format and content of the data, and if the format is correct, sends the data through the cleaning process.
[0762] Step 4:
[0763] Server: The cleaning process removes noise and unnecessary parts from the conversation data and filters out personal information, such as the patient's name and contact details.
[0764] Step 5:
[0765] Server: Organizes the cleaned data and stores it in a database. When storing the data, it categorizes it by conversation topic, person, date and time, etc.
[0766] Step 6:
[0767] Server: Trains a large-scale language model (LLM) using the cleaned conversation data. For example, it trains the model based on the data of psychological counselor C and builds a response generation model specialized for that person.
[0768] Step 7:
[0769] Server: Evaluates the trained model on the test data, adjusts the model parameters if necessary, saves the completed model, and sets up the appropriate API endpoints.
[0770] Step 8:
[0771] Server: Integrates the emotion engine into the AI model. The emotion engine contains algorithms to recognize emotions from user input (text, voice, facial expressions, etc.) and make appropriate adjustments.
[0772] Step 9:
[0773] Terminal: The user launches the application and accesses the login screen. The user enters their credentials and logs in.
[0774] Step 10:
[0775] Server: Checks the user's authentication information, and if authentication is successful, displays the user dashboard, where the user can select the conversation partner they want to simulate (e.g., psychological counselor C).
[0776] Step 11:
[0777] Terminal: Provides an interface for the user to select a simulated conversation partner. The user makes the selection and sends the selection information to the server.
[0778] Step 12:
[0779] Server: To start a simulation based on the user's selection, prepare the corresponding AI model (psychological counselor C's model) and send information to the user's device to optimize performance.
[0780] Step 13:
[0781] Device: Sends text or voice data entered by the user to the server. For example, you might enter, "I've been feeling really stressed lately."
[0782] Step 14:
[0783] Server: Receives user input and passes it to the appropriate AI model. The emotion engine also analyzes this input. The AI model generates a basic response such as "That must be tough," while the emotion engine recognizes emotions such as "stress" and "anxiety" and adjusts the response to "What situations are stressing you out? Tell me about them."
[0784] Step 15:
[0785] Server: Sends the generated response to the user's device.
[0786] Step 16:
[0787] Terminal: Receives the response from the server and displays it in the user interface. The user can then input again if they wish to continue the interaction.
[0788] Step 17:
[0789] Server: After the user's simulation is completed, the server calculates the reward for the data provider based on the simulation time and number of times used, and pays the provider.
[0790] The above is a specific processing flow of the embodiment of this system.
[0791] Example 2
[0792] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0793] Conventional dialogue systems have the following problems. First, the quality of the collected conversation data is inconsistent, resulting in low reliability of the data used for training. Second, they lack the ability to understand the user's emotions and generate responses accordingly, limiting the user experience. Furthermore, the distribution of rewards to data providers is unclear, resulting in a lack of incentive to provide data. These problems must be resolved.
[0794] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0795] In this invention, the server includes means for collecting conversation data, means for cleaning the collected conversation data and storing it in a database, means for generating a large-scale language model trained using the stored conversation data, means for integrating an emotion engine that recognizes user emotions and adjusts responses, and means for calculating and distributing rewards to data providers based on the conversation data used. This makes it possible to generate a reliable large-scale language model using high-quality data and provide appropriate responses according to the user's emotions. Furthermore, transparency of rewards for conversation data providers can be ensured, improving motivation for data provision.
[0796] "Conversational Data" refers to the content of conversations between humans recorded in audio, text, or video format.
[0797] "Means of collection" refers to the mechanism by which user-submitted conversation data is incorporated into the system.
[0798] "Cleaning methods" refers to the process used to remove noise, inappropriate content, and personal information from collected conversation data.
[0799] "Means for storing in a database" refers to the technology used to store the cleaned conversation data as part of a database.
[0800] "Means for generating large-scale language models" refers to a method for using cleaned conversational data to create trained language models based on natural language processing.
[0801] "Means of integrating an emotion engine" refers to the process of incorporating a system into an AI model to recognize emotions from user input data and adjust the model's response based on those emotions.
[0802] "Means for calculating and distributing rewards" refers to a system that calculates and appropriately distributes rewards to data providers based on the amount and frequency of conversation data used.
[0803] "User authentication" refers to the process of verifying a user's identity when logging into a system.
[0804] "Conversation Partner" refers to a virtual or real entity that a user selects within the system with whom to interact.
[0805] "Real-time" means that user input is processed almost instantly and responses are returned immediately.
[0806] The present invention is a system that collects, cleans, and stores conversation data in a database. Furthermore, the system generates a large-scale language model for training using the stored conversation data, and provides an appropriate AI model based on a specific conversation partner after the user is authenticated. The system conducts real-time conversations between the user and the AI model, and the AI model generates responses based on the user's input. The system also includes an emotion engine that recognizes the user's emotions and generates responses adapted to those emotions. The emotion engine can recognize emotions from the user's input voice data and facial expression data. The system also includes a means for calculating and distributing rewards to data providers based on the conversation data used.
[0807] Hardware and software used
[0808] The system uses the following hardware and software:
[0809] Server: A server with high performance data storage and computing power (e.g., AWS EC2 instance)
[0810] Database: A relational database (e.g., Amazon RDS, MySQL) to store conversation data.
[0811] Deep learning frameworks for training language models (e.g., TensorFlow, PyTorch)
[0812] Emotion recognition engine: Software for analyzing voice and facial expression data (e.g., OpenCV, emotionAPI)
[0813] User interface: Web or mobile application for data upload and conversation simulation
[0814] Collecting and cleaning conversation data
[0815] The server provides an interface for conversation data providers to log in and upload conversation data in the form of audio, text, or video. Providers upload the data, which is temporarily stored on the server. The server then cleans the uploaded conversation data, removing noise and filtering inappropriate content and personal information, and stores the cleaned data in a database.
[0816] Generating large-scale language models
[0817] The server uses the cleaned conversation data to train large-scale language models, generates models tailored to specific conversation partners, and sets up API endpoints to make them accessible. Training is done using deep learning frameworks such as TensorFlow and PyTorch.
[0818] Emotion engine integration
[0819] The server integrates an emotion engine into the AI model. The emotion engine recognizes emotions from user input data (text, voice, facial expressions) and adjusts responses based on the results. The emotion recognition algorithm evaluates the user's emotional state in real time and generates appropriate responses.
[0820] User authentication and conversation initiation
[0821] The user launches the application on their device and enters their authentication information on a login screen. If authentication is successful, an interface appears where they can select a conversation partner and begin the dialogue. The user inputs text or voice, and the data is sent to the server. The server passes the received input to an AI model and emotion engine, generating a response in real time. The response is displayed on the device, and the user can continue the conversation.
[0822] Reward distribution
[0823] The server tracks users' conversation data and calculates rewards for data providers based on it. Rewards are calculated based on usage time and frequency and distributed appropriately. This increases transparency for data providers and strengthens their motivation to provide data.
[0824] Specific examples
[0825] For example, consider a case where a user simulates a conversation with a psychological counselor. The user inputs, "I've been feeling very stressed lately," and the message is sent to the server. The emotion engine analyzes this input and recognizes the user's emotion as "stress." The AI model generates a basic response, "That must be tough," which the emotion engine then refines by saying, "What situations are causing you stress? Tell me about them." This response is ultimately returned to the user.
[0826] Prompt Sentence Examples
[0827] Example prompt: "Tell me about the latest robotics technology."
[0828] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0829] Step 1:
[0830] Conversation data collection
[0831] Input: Audio, text, or video files provided by the conversation data provider
[0832] Server: A conversation data provider logs in to the system and uploads conversation data through the data upload interface. When the provider selects the data and clicks the upload button, the data is sent to the server and temporarily stored.
[0833] Output: Temporarily saved conversation data
[0834] Specific behavior:
[0835] The user clicks the "Upload conversation data" button.
[0836] The selected file is sent to the server and saved in a temporary folder.
[0837] Step 2:
[0838] Data Preprocessing
[0839] Input: Temporarily saved conversation data
[0840] Server: The uploaded conversation data is cleaned. This process includes removing noise and filtering inappropriate content and personal information. The cleaned data is then stored in a database.
[0841] Output: Cleaned conversation data
[0842] Specific behavior:
[0843] The server adds the uploaded data to a processing queue.
[0844] A cleaning algorithm steps through the data in the queue and applies a noise reduction filter.
[0845] Apply a set of personal information filtering rules to automatically remove inappropriate content.
[0846] Store clean data in the database.
[0847] Step 3:
[0848] AI model generation
[0849] Input: Cleaned conversation data
[0850] Server: Trains large-scale language models using cleaned conversation data, generates models tailored to specific conversation partners, and sets up API endpoints to make them accessible.
[0851] Output: A trained large-scale language model
[0852] Specific behavior:
[0853] The server retrieves the cleaned data from the database.
[0854] Train the model using a deep learning framework (e.g., PyTorch, TensorFlow).
[0855] Save the parameters of the model once it has been trained.
[0856] Create an API endpoint and deploy the model.
[0857] Step 4:
[0858] Emotion engine integration
[0859] Input: A trained large-scale language model
[0860] Server: Integrates an emotion engine into the AI model. The emotion engine contains algorithms for recognizing emotions from user input data (text, voice, facial expressions) and adjusting responses based on the results.
[0861] Output: AI model with integrated emotion engine
[0862] Specific behavior:
[0863] Emotion recognition algorithms perform text and voice analysis.
[0864] Assessing the user's emotional state (e.g., happy, sad, angry) in real time.
[0865] The basic responses generated by the AI model are fine-tuned based on the emotion recognition results to generate the optimal response.
[0866] Step 5:
[0867] User authentication and selection
[0868] Input: Authentication information (email address, password)
[0869] On the device: The user launches the application and accesses the login screen. The user enters their email address and password for authentication. If successful, the conversation partner selection screen is displayed.
[0870] Output: List of conversation partners after successful authentication
[0871] Specific behavior:
[0872] The user launches the application and enters their login ID and password.
[0873] The server verifies the authentication information and, if it matches, starts the session.
[0874] After successful login, a list of conversation partners will be displayed on your device.
[0875] Step 6:
[0876] Conversation Simulation
[0877] Input: User text or voice input
[0878] Terminal: The user selects a conversation partner through the system's interface and starts a conversation. When the user inputs text or voice, the data is sent to the server.
[0879] Server: Receives user input and passes it to the appropriate AI model and emotion engine. The AI model generates a basic response, and the emotion engine adjusts the response based on the user's emotion.
[0880] Output: Adjusted AI response
[0881] Specific behavior:
[0882] The user selects "Psychological Counselor C" and clicks the "Start Dialogue" button.
[0883] The user types "I've been feeling really stressed lately" into the input field and submits.
[0884] The server receives this input and sends it to the AI model.
[0885] The model generates the response "That's tough," which the emotion engine refines to "What situations are stressing you out? Tell me about them."
[0886] The response is returned to the terminal and displayed to the user.
[0887] Step 7:
[0888] Reward distribution
[0889] Input: conversation session data (usage time, usage frequency)
[0890] Server: Tracks user conversation records and calculates rewards for data providers based on them. Calculates rewards based on the time and frequency of use and pays providers.
[0891] Output: Calculated reward points and payment processing
[0892] Specific behavior:
[0893] The server records data for each conversation session.
[0894] A reward calculation algorithm calculates points based on the duration and frequency of use per session.
[0895] Rewards are transferred to conversation data providers through a payment processing service.
[0896] (Application example 2)
[0897] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0898] Customer service in traditional brick-and-mortar stores is heavily dependent on the experience and skills of store staff, making it difficult to maintain consistent service quality. Furthermore, in situations where flexible responses based on customer emotions are required, it is difficult for humans alone to recognize and respond 100% accurately. Furthermore, the appropriate management and effective use of collected conversation data is an issue.
[0899] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting conversation data, means for cleaning the collected conversation data and storing it in a database, means for generating a large-scale language model to be trained using the stored conversation data, means for providing an appropriate AI model based on a specific conversation partner selected by an authenticated user, means for conducting a real-time conversation between the user and the AI model and for the AI model to generate a response based on the user's input, means for generating a response based on the user's emotions using an emotion recognition engine, means for analyzing collected audio and video data using smart glasses and providing appropriate responses to store clerks in real time, and means for calculating and distributing rewards to data providers based on the used conversation data. This enables real-time responses that take customer emotions into consideration and uniform service quality.
[0900] "Conversation Data" means records of audio, text, and video communications between users or between users and the system.
[0901] "Cleaning" is the process of removing noise and organizing data while removing inappropriate content and personal information.
[0902] A "database" is a collection of information that stores collected data in an organized manner and makes it easily searchable and accessible.
[0903] A "large-scale language model" is an artificial intelligence model trained using vast amounts of conversational data to generate natural-sounding language responses based on user input.
[0904] An "AI model" is an artificial intelligence algorithm and its implementation designed to address a specific task or application.
[0905] An "emotion recognition engine" analyzes emotions from a user's voice and video data and adjusts responses based on those emotions.
[0906] "Smart glasses" are a type of wearable device used to acquire and display visual and audio information, and have the ability to present analysis results to users in real time.
[0907] "Reward" means compensation distributed to data providers based on the use of the conversation data they provide.
[0908] "Data providers" are individuals or organizations that upload conversation data to the system.
[0909] The present invention is a system that aims to improve the quality and efficiency of customer support in brick-and-mortar stores, and includes means for collecting and cleaning conversation data, generating large-scale language models, recognizing emotions, and distributing rewards. The following describes in detail the embodiments of the present invention.
[0910] System Program
[0911] 1. Collecting conversation data
[0912] The server collects conversation data through smart glasses worn by store staff. The smart glasses are equipped with a microphone to collect voice data and a camera to capture facial expression data. The collected data is both audio and video.
[0913] 2. Data Preprocessing
[0914] The server receives the raw conversation data sent by the smart glasses and cleans it. The cleaning process includes removing noise and filtering inappropriate content and personal information, improving the quality of the data before storing it in a database.
[0915] 3. Generating large-scale language models
[0916] The server uses the cleaned conversation data to train a large-scale language model, which is trained specifically for specific conversational partners (in this case, the store clerk and the customer) and made accessible through an API, using Python and AI frameworks such as TensorFlow.
[0917] 4. Emotion Recognition Integration
[0918] The server integrates a large-scale language model with an emotion recognition engine, which uses OpenCV and TensorFlow to analyze user emotions from audio and video data. The analysis results are reflected in the generated response.
[0919] 5. User Authentication and Response Generation
[0920] The terminal (smart glasses) performs personal authentication when the store clerk logs in. After logging in, when the store clerk interacts with the customer, voice input and video data are sent to the server. The server receives this data and passes it to an AI model and emotion recognition engine to generate an appropriate response.
[0921] 6. Real-time response
[0922] The server sends the generated response to the smart glasses in real time and displays it to the store clerk. This allows the store clerk to immediately provide an appropriate response to the customer. For example, if a customer asks, "I've been feeling stressed lately," the smart glasses will respond, "That's tough. What situations are making you feel stressed?"
[0923] 7. Reward Distribution
[0924] The server calculates and distributes rewards to data providers (store clerks and stores) based on the collected conversation data. Rewards are determined based on the frequency and quality of data usage.
[0925] Hardware and software used
[0926] Hardware: Smart glasses (e.g., Google Glass Enterprise Edition 2)
[0927] Server software: AWS Lambda, S3, EC2
[0928] Preprocessing and training: Python scripts, TensorFlow
[0929] Emotion recognition: OpenCV, TensorFlow
[0930] Authentication system: OAuth 2.0, Firebase Authentication
[0931] Examples of prompt statements
[0932] For example, if a customer asks, "I've been feeling stressed lately, especially at work...what should I do?" the system might respond with:
[0933] "That's tough. What situations are stressing you out?"
[0934] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0935] Step 1:
[0936] A user (a store clerk) wears smart glasses and serves customers in a store. The smart glasses collect audio and video data in real time and send it to a server. The input is conversation data (audio and video) between the customer and the store clerk, which is then sent to the server.
[0937] Step 2:
[0938] The server receives the transmitted conversation data and performs cleaning processes such as noise removal, filtering of inappropriate content and personal information. The data is then cleaned and stored in a database. The input is raw conversation data, and the output is cleaned data.
[0939] Step 3:
[0940] The server uses the cleaned conversation data to train a large-scale language model. It uses an AI framework such as TensorFlow to generate a model specialized for a specific conversation partner. The input is the cleaned conversation data, and the output is the trained large-scale language model.
[0941] Step 4:
[0942] The server integrates an emotion recognition engine with the trained large-scale language model. This engine uses OpenCV and TensorFlow to recognize user emotions from audio and video data. The input is audio and video data, and the output is recognized emotional information and response adjustments based on it.
[0943] Step 5:
[0944] The smart glasses at the terminal authenticate the store clerk by logging in to the system. During this process, the user's authentication information is sent to the server and verified. The input is the store clerk's authentication information, and the output is the authentication success or failure status.
[0945] Step 6:
[0946] The server receives real-time audio and video data when an authenticated store associate interacts with a customer and passes it to an AI model and emotion recognition engine. The AI model and emotion recognition engine then generate an appropriate response based on the user's input data. The input is real-time conversation data with the customer, and the output is the generated appropriate response.
[0947] Step 7:
[0948] The server sends the generated response to the smart glasses in real time and presents it to the store clerk. The store clerk responds to the customer based on the response. The input is the generated response, and the output is the response displayed on the smart glasses.
[0949] Step 8:
[0950] The server tracks all conversation sessions and calculates and distributes rewards to data providers based on the conversation data used. The input is the conversation session data, and the output is the reward calculation and distribution results.
[0951] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0952] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0953] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0954] [Third embodiment]
[0955] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0956] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0957] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0958] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0959] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0960] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0961] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0962] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0963] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0964] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0965] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0966] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0967] The present invention provides a system that collects, cleans, and stores conversation data in a database, generates a large-scale language model for training using the stored conversation data, and provides an appropriate AI model based on a specific conversation partner after a user is authenticated. The system also conducts real-time conversations between the user and the AI model, and the AI model generates responses based on the user's input. The system also includes a means for calculating and distributing rewards to data providers based on the conversation data used.
[0968] Program processing
[0969] 1. Collecting conversation data
[0970] Server: Provides an interface for conversation data providers (celebrities and professionals) to upload their conversation data after logging in. When providers upload audio, text, or video data, the data is temporarily stored.
[0971] 2. Data Preprocessing
[0972] Server: Analyzes the received conversation data, removes noise and unnecessary parts, and filters out personal information. The preprocessed data is saved in the database as cleaned data.
[0973] 3. AI model generation
[0974] Server: Uses the cleaned data to train large-scale language models, which are then tuned to specific conversational partners and set up API endpoints to make them accessible.
[0975] 4. User Authentication and Selection
[0976] Terminal: Provides an authentication interface for users to access the system through the application and log in. Once authenticated, users can select a specific conversation partner from a list.
[0977] 5. Conversation Simulation
[0978] Terminal: The terminal displays an input form on the interface to start a conversation with a conversation partner of the user's choice. The user enters the information and sends it to the server.
[0979] Server: Receives user input and passes it to the appropriate AI model, which generates a response and sends it back to the device.
[0980] Terminal: Displays responses from the server and provides an environment in which the user can continue the dialogue.
[0981] 6. Reward Distribution
[0982] Server: Tracks the user's conversation time and frequency of use, and calculates the reward for the conversation data provider. Once the reward is calculated, payment is processed based on the provider's account information.
[0983] Specific examples
[0984] For example, consider the case where a famous actor B provides audio data about his or her stage experience. This data is collected, cleaned to remove personal information, and then stored in a database. Next, a large-scale language model is trained based on Actor B's audio data to generate an AI model that can converse with Actor B.
[0985] A user logs into this system and selects "Simulated dialogue with actor B." When the user inputs "What was your most recent stage challenge?" to actor B, the input is sent to the server, and the AI model generates a response such as "My most recent stage challenge was trying out a new role," and sends it back to the user.
[0986] When the conversation ends, the server calculates the reward for Actor B based on the time and frequency of use and transfers it to the system. In this way, users can advance their language learning through a realistic dialogue experience, and the provider of the conversation data can also earn rewards.
[0987] This system not only allows language learners to stay motivated and improve their skills effectively, but also allows conversation data providers to earn revenue in new ways.
[0988] The processing flow will be explained below.
[0989] Understood. Below is the process flow broken down into specific steps.
[0990] Step 1:
[0991] Server: Provides an authentication interface for conversation data providers (e.g., famous actor B) to log in. The conversation data provider logs in by entering the correct authentication information.
[0992] Step 2:
[0993] Server: After successful authentication, the server displays a conversation data upload interface to the conversation data provider, where the provider can select and upload conversation data in audio, text, or video format.
[0994] Step 3:
[0995] Server: Temporarily stores the uploaded data, checks the format and content of the data, and if the format is correct, sends the data to the cleaning process.
[0996] Step 4:
[0997] Server: The cleaning process removes noise and unnecessary parts from the conversation data and filters out personal information, such as the patient's name and contact details.
[0998] Step 5:
[0999] Server: Organizes the cleaned data and stores it in a database. When storing the data, it categorizes it by conversation topic, person, date and time, etc.
[1000] Step 6:
[1001] Server: Uses the cleaned conversation data to train a large-scale language model (LLM). For example, it trains the model based on actor B's data and builds a response generation model specialized for that person.
[1002] Step 7:
[1003] Server: Evaluates the trained model on the test data, adjusts the model parameters if necessary, saves the completed model, and sets up the appropriate API endpoints.
[1004] Step 8:
[1005] Terminal: The user launches the application and accesses the login screen. The user enters their credentials and logs in.
[1006] Step 9:
[1007] Server: Checks the user's authentication information, and if authentication is successful, displays the user dashboard, where the user can select the conversation partner they want to simulate (e.g., Actor B).
[1008] Step 10:
[1009] Terminal: Displays an interface for the user to select a simulated conversation partner. The user makes the selection and sends the selection information to the server.
[1010] Step 11:
[1011] Server: Based on the user's selection, prepares the corresponding AI model (e.g., actor B's model) and sends interface information to the terminal to start the conversation.
[1012] Step 12:
[1013] Terminal: Sends a text message typed by the user to the server. For example, "What's your latest stage challenge?"
[1014] Step 13:
[1015] Server: Receives user input and passes it to the appropriate AI model, which then generates a response, for example, "My most recent stage challenge was taking on a new role."
[1016] Step 14:
[1017] Server: Sends the generated response to the user's device.
[1018] Step 15:
[1019] Terminal: Receives the response from the server and displays it in the user interface. If the user wants to continue the interaction, they can input again.
[1020] Step 16:
[1021] Server: After the user's simulation is completed, the server calculates the reward for the conversation data provider based on the simulation time and number of times used.
[1022] Step 17:
[1023] Server: Carries out the payment procedures to distribute the calculated reward to the conversation data provider. The reward is transferred based on the provider's account information.
[1024] The above is the specific processing flow of this system.
[1025] Example 1
[1026] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1027] Conventional conversation data collection and processing systems have insufficient quality control of collected data, particularly due to the inclusion of noise and personal information, which limits its use. Furthermore, when training AI models using conversation data, it is difficult to adjust them to specific conversation partners, making it difficult to provide real-time responses to users. Furthermore, reward distribution to data providers is sometimes not transparent and efficient, resulting in a lack of incentives for providers.
[1028] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1029] In this invention, the server includes a means for collecting conversation data, a means for cleaning the collected conversation data and storing it in a database, and a means for generating a model to be trained using the stored conversation data. This enables appropriate preprocessing of the data and the generation of an AI model tailored to a specific conversation partner using high-quality training data. The server also includes a means for selecting a specific conversation partner after the user is authenticated, a means for conducting a conversation between the user and the model and for the model to generate a response based on the user's input, and a means for calculating and distributing rewards to data providers based on the used conversation data. This allows users to receive appropriate responses in real time and distributes rewards fairly and quickly to data providers, improving the efficiency and satisfaction of the entire system.
[1030] "Conversation data" refers to audio data, text data, or video data provided to the system by users or conversation data providers.
[1031] "Cleaning" is the process of removing noise, inappropriate content, personal information, etc. from collected conversation data to improve data quality before storing it in a database.
[1032] A "database" is a storage within the system that efficiently and securely stores cleaned conversation data for later use as training data.
[1033] "Training" is the process of using cleaned conversational data to train an AI model and improve its pattern recognition capabilities and prediction accuracy.
[1034] A "model" is a collection of AI structures and algorithms trained on conversational data that are used to generate appropriate responses to user input.
[1035] A "specific conversation partner" is a virtual conversation partner for a dialogue that is selected by the user based on various categories and profiles provided on the system.
[1036] "Real-time" means that there is a very short delay between when a user makes an input and when the AI model returns a response, and refers to a situation in which the interaction takes place immediately.
[1037] A "response" is a response, such as text or voice, that an AI model generates based on user input.
[1038] "Remuneration" is the compensation calculated and paid to the conversation data provider based on the degree and frequency of use of the provided data.
[1039] "Distribution" is the process of appropriately paying the calculated reward based on the registered information of the conversation data provider.
[1040] "Interface" refers to the screens, input forms, and operating means used by conversation data providers when uploading data or by users when accessing and operating the system.
[1041] To implement this invention, the server, terminal, and user each play specific roles. The main hardware used includes a database server, web server, and user terminal (PC or smartphone), while the software required includes a framework for training large-scale language models (e.g., TensorFlow, PyTorch), database software (e.g., MySQL, PostgreSQL), and a web service API.
[1042] The server provides a means for collecting conversation data. Specifically, it provides a user interface for conversation data providers to log in and upload audio data, text data, and video data. The device supports a means for providers to upload data, allowing them to select and send data in various formats (e.g., .mp3, .txt, .mp4). Uploaded data is temporarily stored on the server.
[1043] The server analyzes the collected conversation data, removes noise and unnecessary parts, and filters out personal information to clean it. The cleaned data is then stored in a database, which later serves as the foundation for training AI models.
[1044] The server then uses the cleaned conversation data to train a large-scale language model using a large-scale language model framework, processing the data in batches to learn the model weights, resulting in an AI model that corresponds to a specific conversation partner. The model is then made accessible through an API endpoint.
[1045] To access the system, a user uses a terminal to enter a username and password into a login interface. After successful authentication, the server provides the user with a list of specific available conversation partners from which the user can select a conversation partner.
[1046] The device then displays an interface for the user to converse with the selected conversation partner. The user enters a question or message into the input form and clicks the send button. For example, a question such as "What was your most recent stage challenge?" is sent to the server.
[1047] The server receives user input, passes it to the trained AI model, and generates a response, which is then sent back to the device, where it is displayed to the user, allowing the user to continuously engage in real-time dialogue with the AI model.
[1048] Finally, the server tracks the duration and frequency of users' conversation data usage and calculates rewards for conversation data providers based on that data. The calculated rewards are distributed based on the providers' registered account information.
[1049] As a concrete example, consider a case where a user selects a dialogue simulation with actor B and asks, "What was the most moving moment on stage?" This question is transmitted to the AI model via the server, and the AI model generates a response such as, "The most moving moment on stage was when the audience shed tears," and sends it back to the user.
[1050] This system allows users to advance their language learning through realistic conversational experiences, and enables conversation data providers to earn revenue in new ways.
[1051] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1052] Step 1:
[1053] The server verifies the login of the conversation data provider and, after successful authentication, provides an interface for data upload.
[1054] Input: Provider username and password.
[1055] Specific operation: Query the database to verify the provider's authentication information, and if authentication is successful, start a login session.
[1056] Output: Login session token and data upload interface.
[1057] Step 2:
[1058] Using the terminal, the conversation data provider selects audio data, text data, or video data and clicks the upload button.
[1059] Input: Data files selected by the provider (e.g. .mp3, .txt, .mp4).
[1060] Specific operation: A file selection dialog is displayed, and after the provider selects a file, they press the upload button.
[1061] Output: The selected data file is sent to the server.
[1062] Step 3:
[1063] The server stores the uploaded data in temporary storage and checks the integrity of the data.
[1064] Input: Uploaded data file.
[1065] What it does: Saves the file to temporary storage and verifies the data integrity by checking the file format and the first few bytes of the data.
[1066] Output: Data file with integrity checked.
[1067] Step 4:
[1068] The server analyzes the received conversation data, removes noise and unnecessary parts, and filters out personal information to clean it.
[1069] Input: Data file whose integrity has been checked.
[1070] Specific operation: For audio data, a noise reduction algorithm is applied, and for text data, filtering is performed to mask specific keywords.
[1071] Output: Cleaned data.
[1072] Step 5:
[1073] The server organizes and stores the cleaned data in a database.
[1074] Input: Cleaned data.
[1075] What it does: Inserts data and creates indexes according to the appropriate schema in the database.
[1076] Output: Cleaned data stored in a database.
[1077] Step 6:
[1078] The server uses the cleaned conversational data to train a large-scale language model.
[1079] Input: Cleaned data in the database.
[1080] Specific behavior: Use a large-scale language model framework (e.g., TensorFlow, PyTorch) to process data in batches and update the model weights.
[1081] Output: A trained AI model.
[1082] Step 7:
[1083] The server tunes the trained model to correspond to specific conversation partners and sets up an API endpoint to make it accessible externally.
[1084] Input: A trained AI model.
[1085] Specific actions: Adjust model parameters, configure API endpoints, and deploy.
[1086] Output: An accessible API endpoint for the AI model.
[1087] Step 8:
[1088] The device presents the user with a login interface, prompting them to enter their username and password, and upon successful authentication, providing an interface for selecting a specific conversation partner.
[1089] Input: Username and Password.
[1090] Specific behavior: The user enters information into the login form and clicks the login button.
[1091] Output: Conversation partner selection interface for authenticated user.
[1092] Step 9:
[1093] The terminal displays a screen for starting a conversation based on the conversation partner selected by the user.
[1094] Input: User's selected conversation partner information.
[1095] Specific behavior: Displays a conversation interface and provides an input form.
[1096] Output: A form for the user to enter.
[1097] Step 10:
[1098] The user enters a question or message into the input form and clicks the send button.
[1099] Input: User question or message (e.g., "What was your most recent stage challenge?").
[1100] Specific behavior: The user enters text and clicks the submit button.
[1101] Output: The entered question or message is sent to the server.
[1102] Step 11:
[1103] The server receives user input and passes it to a trained AI model to generate a response.
[1104] Input: The user's typed question or message.
[1105] What it does: Tokenizes the user's message and passes it to the AI model for processing.
[1106] Output: The generated response of the AI model.
[1107] Step 12:
[1108] The terminal displays responses from the server to the user, providing an interface that allows the conversation to continue.
[1109] Input: The response sent back from the server.
[1110] Specific behavior: Displays the response text and persists the interface so the user can enter it again.
[1111] Output: The response displayed to the user and an input form for further interaction.
[1112] Step 13:
[1113] The server tracks the time and frequency of users' conversation data usage and calculates rewards for conversation data providers based on that data.
[1114] Input: User conversation time, usage frequency data.
[1115] Specific operation: Analyzes conversation logs and calculates rewards based on an algorithm.
[1116] Output: Calculated reward data.
[1117] Step 14:
[1118] The server distributes the calculated reward based on the provider's registered account information.
[1119] Input: Remuneration data and provider account information.
[1120] Specific operation: The reward is transferred to the provider's account.
[1121] Output: Rewards deposited into the provider's account.
[1122] The above are the specific processing steps of the program in this system.
[1123] (Application example 1)
[1124] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1125] The present invention relates to a technology and system that provides dialogue simulations to meet the needs of users who want to have real-time dialogues with specific conversation partners. It also aims to solve the problem of simultaneously providing a reward distribution mechanism for dialogue data providers. This allows users to enhance their learning by engaging in dialogues that include specialized knowledge, and data providers to gain new revenue sources.
[1126] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1127] In this invention, the server includes means for collecting conversation data, means for cleaning the collected conversation data and storing it in a database, means for generating a large-scale language model for training using the stored conversation data, means for providing an appropriate AI model based on a specific conversation partner selected by an authenticated user, means for conducting a conversation between the user and the AI model in real time and for the AI model to generate a response based on the user's input, means for calculating and distributing rewards to data providers based on the used conversation data, means for providing an interface for the user to start a dialogue simulation, and means for tracking the user's conversation data and transferring rewards to dialogue data providers through the system. This allows users to simulate a dialogue with a specific conversation partner in real time and effectively learn and exchange information, while also allowing dialogue data providers to receive fair rewards.
[1128] "Conversation Data" means information in the form of audio, text, or video exchanged between Users or between Users and the System.
[1129] "Cleaning" refers to the process of removing noise, inappropriate content, and personal information from collected conversation data.
[1130] "Database" refers to a system that accumulates and manages cleaned conversation data.
[1131] A "large-scale language model" refers to a natural language processing model trained on a massive amount of conversational data.
[1132] "AI model" refers to a large-scale language model tuned based on a specific conversation partner.
[1133] "Real-time" refers to a situation where there is almost no delay and processing or response is immediate.
[1134] "Conversation Partner" refers to a specific expert, celebrity, or AI model thereof with whom the user wishes to have a conversation.
[1135] "Interface" refers to the screen and input means that users use to operate the system.
[1136] "Tracking" refers to the act of following and recording user behavior and conversation data.
[1137] "Reward" refers to money or other consideration paid by the system to the dialogue data provider.
[1138] The present invention is a system that allows users to have real-time conversations with specific conversation partners. The system provides a series of functions including collection, cleaning, and storage of conversation data, generation of AI models, user authentication, real-time conversation execution, and reward distribution.
[1139] System program generation
[1140] Data collection and cleaning
[1141] The server provides an interface for conversation data providers (such as experts and celebrities) to upload audio, text, and video data after they log in. This data is temporarily stored and analyzed to filter out noise, unnecessary parts, and personal information. Once preprocessing is complete, the data is stored in a database as cleaned data.
[1142] Generating AI models
[1143] The server uses the cleaned data to train large-scale language models, which are then tuned to specific conversational partners and set up API endpoints to make them accessible.
[1144] User authentication and conversation simulation
[1145] The terminal (smartphone) provides an authentication interface for users to access and log in to the system. Once authentication is complete, the user is presented with an interface to select a specific conversation partner and begin the dialogue simulation. The user's input is sent to the server, and the corresponding AI model generates a response and sends it back to the terminal. The server performs this process in real time, providing an environment in which the user can continue the dialogue.
[1146] Reward distribution
[1147] The server tracks the user's conversation time and frequency of use, calculates the reward for the conversation data provider, and then processes the payment based on the provider's account information.
[1148] Hardware and software used
[1149] Hardware: A server with a high-performance CPU and GPU, and a smartphone (iOS or Android device)
[1150] Software: Applications running on edge devices (Swift or Kotlin), Python scripts for data analysis and filtering, database management system (MySQL), large-scale language model (OpenAI GPT), API server (Flask or Django)
[1151] Specific examples
[1152] For example, if a user logs in to the system and selects a simulated conversation with a specific expert, the following process will occur: When the user types, "Tell me about your new business idea," the server receives this input and passes it to an AI model trained as a specific expert. The AI model generates a response: "The key to a new business idea is market research and finding a niche." The generated response is sent back to the terminal, and the user can use it to ask further questions.
[1153] Prompt Sentence Examples
[1154] "If a user asks, 'Tell me about a new business idea,' generate a response from an AI trained as a business expert."
[1155] This allows users to enhance their learning by engaging in conversations that include specialized knowledge, and provides conversation data providers with a new source of revenue.
[1156] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1157] Program processing steps
[1158] Step 1:
[1159] Conversation data collection
[1160] Server: When a conversation data provider logs in, it provides an interface for uploading voice data, text data, and video data. Once the provider uploads the data, it is temporarily stored.
[1161] Input: Conversation data (audio, text, video)
[1162] Output: Temporarily saved conversation data
[1163] Specific operation: After authenticating the provider, the server stores the uploaded data in a specific directory on the server.
[1164] Step 2:
[1165] Data Preprocessing
[1166] Server: Analyzes the collected conversation data, filters out noise, unnecessary parts, and personal information, and stores the cleaned data in a database.
[1167] Input: Temporarily saved conversation data
[1168] Output: Cleaned conversation data
[1169] What it does: It uses a Python script to denoise the audio data, sanitize the text, and filter out personal information, then stores it in a MySQL database.
[1170] Step 3:
[1171] Generating AI models
[1172] Server: Trains large-scale language models using cleaned data, tunes AI models based on specific conversational partners, and sets up API endpoints for access.
[1173] Input: Cleaned conversation data
[1174] Output: A tuned AI model
[1175] What it does: Train and tune large-scale language models using Python and TensorFlow. Set up API endpoints using Flask or Django.
[1176] Step 4:
[1177] User authentication and selection
[1178] Terminal (smartphone): Provides an authentication interface for users to access the system through the app and log in. Once authenticated, users can select a specific conversation partner.
[1179] Input: User login information, conversation partner selection
[1180] Output: Preparation for starting the conversation simulation
[1181] Specific operation: The login screen is displayed on the device using Swift or Kotlin, and the user's input information is sent to the server for authentication. After authentication, a list of conversation partners is displayed.
[1182] Step 5:
[1183] Conversation Simulation
[1184] Terminal (smartphone): Displays an interface that allows the user to start a conversation with a selected conversation partner and sends the user's input to the server.
[1185] Server: Passes user input to the AI model, generates a response, and sends it back to the device, which displays the response to the user.
[1186] Input: User's spoken input
[1187] Output: Response by the AI model
[1188] Specific operation: Receives user input and sends it to the server, which passes the input to the AI model and returns the generated response to the mobile device, which displays the response.
[1189] Step 6:
[1190] Reward distribution
[1191] Server: Tracks user conversation time and frequency of use, calculates rewards for conversation data providers, and processes payments.
[1192] Input: Conversation transcript and usage data
[1193] Output: Reward calculation and payment
[1194] Specific operation: A Python script aggregates conversation time and frequency of use, calculates the reward for the provider, and processes the payment to the provider's registered account.
[1195] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1196] The present invention provides a system that collects, cleans, and stores conversation data in a database. It also generates a large-scale language model for training using the stored conversation data, and provides an appropriate AI model based on a specific conversation partner after a user is authenticated. It also conducts real-time conversations between the user and the AI model, and the AI model generates responses based on the user's input. It also includes an emotion engine that recognizes the user's emotions and generates responses adapted to those emotions. This emotion engine can recognize emotions from the user's input voice data and facial expression data. It also includes a means for calculating and distributing rewards to data providers based on the conversation data used.
[1197] Program processing
[1198] 1. Collecting conversation data
[1199] Server: Conversation data providers log in to the system and upload data in the form of audio, text, or video through the conversation data upload interface. This data is temporarily stored on the server.
[1200] 2. Data Preprocessing
[1201] Server: The uploaded conversation data is cleaned. This process includes removing noise and filtering inappropriate content and personal information. The cleaned data is then stored in a database.
[1202] 3. AI model generation
[1203] Server: Trains large-scale language models using cleaned conversation data, trains models specific to specific conversation partners, and sets up API endpoints for access.
[1204] 4. Emotion engine integration
[1205] Server: Integrates an emotion engine into the AI model. The emotion engine contains algorithms for recognizing emotions from user input (text, voice, facial expressions, etc.) and adjusting responses based on the recognized emotions.
[1206] 5. User Authentication and Selection
[1207] Terminal: The user launches the application and accesses the login screen. The user enters their credentials and logs in.
[1208] 6. Conversation Simulation
[1209] Terminal: The user selects a conversation partner through the system's interface and starts a dialogue. The terminal transmits the text and voice data entered by the user to the server.
[1210] Server: Receives user input and passes it to the appropriate AI model and emotion engine. The AI model generates a basic response, and the emotion engine tailors the response based on the user's emotion. For example, if a user makes an input that sounds like they're in distress, the emotion engine will recognize that and generate a response like, "You seem distressed. Is there anything I can help you with?"
[1211] Terminal: displays the adjusted response to the user and allows the user to continue typing.
[1212] 7. Reward Distribution
[1213] Server: Tracks user conversation session data and calculates rewards for data providers based on the time and frequency of use, and pays providers.
[1214] Specific examples
[1215] For example, consider the case where psychological counselor C provides audio data from a session. Counselor C's data is collected, cleaned, and stored in a database. Next, a large-scale language model is trained based on psychological counselor C's data to build an AI model that generates specific responses.
[1216] A user logs into the system and selects "Simulated conversation with psychological counselor C." When the user enters "I've been feeling very stressed lately," the message is sent to the server. The emotion engine analyzes this input and recognizes the user's emotions as "stress" or "anxiety." The AI model generates a basic response of "That must be tough," which the emotion engine refines to "What situations are causing you stress? Tell me about them." This response is ultimately returned to the user.
[1217] This system allows users to have a realistic conversational experience, and allows them to receive psychological support while learning a language. Conversation data providers are also rewarded based on the use of their data.
[1218] The processing flow will be explained below.
[1219] Understood. Below is the process flow broken down into specific steps.
[1220] Step 1:
[1221] Server: Provides an authentication interface for conversation data providers to log in. The conversation data provider logs in by entering the correct authentication information.
[1222] Step 2:
[1223] Server: After successful authentication, the server displays an interface for uploading conversation data to the data provider, allowing the provider to select and upload data in audio, text, or video format.
[1224] Step 3:
[1225] Server: Temporarily stores the uploaded data, checks the format and content of the data, and if the format is correct, sends the data through the cleaning process.
[1226] Step 4:
[1227] Server: The cleaning process removes noise and unnecessary parts from the conversation data and filters out personal information, such as the patient's name and contact details.
[1228] Step 5:
[1229] Server: Organizes the cleaned data and stores it in a database. When storing the data, it categorizes it by conversation topic, person, date and time, etc.
[1230] Step 6:
[1231] Server: Trains a large-scale language model (LLM) using the cleaned conversation data. For example, it trains the model based on the data of psychological counselor C and builds a response generation model specialized for that person.
[1232] Step 7:
[1233] Server: Evaluates the trained model on the test data, adjusts the model parameters if necessary, saves the completed model, and sets up the appropriate API endpoints.
[1234] Step 8:
[1235] Server: Integrates the emotion engine into the AI model. The emotion engine contains algorithms to recognize emotions from user input (text, voice, facial expressions, etc.) and make appropriate adjustments.
[1236] Step 9:
[1237] Terminal: The user launches the application and accesses the login screen. The user enters their credentials and logs in.
[1238] Step 10:
[1239] Server: Checks the user's authentication information, and if authentication is successful, displays the user dashboard, where the user can select the conversation partner they want to simulate (e.g., psychological counselor C).
[1240] Step 11:
[1241] Terminal: Provides an interface for the user to select a simulated conversation partner. The user makes the selection and sends the selection information to the server.
[1242] Step 12:
[1243] Server: To start a simulation based on the user's selection, prepare the corresponding AI model (psychological counselor C's model) and send information to the user's device to optimize performance.
[1244] Step 13:
[1245] Device: Sends text or voice data entered by the user to the server. For example, you might enter, "I've been feeling really stressed lately."
[1246] Step 14:
[1247] Server: Receives user input and passes it to the appropriate AI model. The emotion engine also analyzes this input. The AI model generates a basic response such as "That must be tough," while the emotion engine recognizes emotions such as "stress" and "anxiety" and adjusts the response to "What situations are stressing you out? Tell me about them."
[1248] Step 15:
[1249] Server: Sends the generated response to the user's device.
[1250] Step 16:
[1251] Terminal: Receives the response from the server and displays it in the user interface. The user can then input again if they wish to continue the interaction.
[1252] Step 17:
[1253] Server: After the user's simulation is completed, the server calculates the reward for the data provider based on the simulation time and number of times used, and pays the provider.
[1254] The above is a specific processing flow of the embodiment of this system.
[1255] Example 2
[1256] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1257] Conventional dialogue systems have the following problems. First, the quality of the collected conversation data is inconsistent, resulting in low reliability of the data used for training. Second, they lack the ability to understand the user's emotions and generate responses accordingly, limiting the user experience. Furthermore, the distribution of rewards to data providers is unclear, resulting in a lack of incentive to provide data. These problems must be resolved.
[1258] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1259] In this invention, the server includes means for collecting conversation data, means for cleaning the collected conversation data and storing it in a database, means for generating a large-scale language model trained using the stored conversation data, means for integrating an emotion engine that recognizes user emotions and adjusts responses, and means for calculating and distributing rewards to data providers based on the conversation data used. This makes it possible to generate a reliable large-scale language model using high-quality data and provide appropriate responses according to the user's emotions. Furthermore, transparency of rewards for conversation data providers can be ensured, improving motivation for data provision.
[1260] "Conversational Data" refers to the content of conversations between humans recorded in audio, text, or video format.
[1261] "Means of collection" refers to the mechanism by which user-submitted conversation data is incorporated into the system.
[1262] "Cleaning methods" refers to the process used to remove noise, inappropriate content, and personal information from collected conversation data.
[1263] "Means for storing in a database" refers to the technology used to store the cleaned conversation data as part of a database.
[1264] "Means for generating large-scale language models" refers to a method for using cleaned conversational data to create trained language models based on natural language processing.
[1265] "Means of integrating an emotion engine" refers to the process of incorporating a system into an AI model to recognize emotions from user input data and adjust the model's response based on those emotions.
[1266] "Means for calculating and distributing rewards" refers to a system that calculates and appropriately distributes rewards to data providers based on the amount and frequency of conversation data used.
[1267] "User authentication" refers to the process of verifying a user's identity when logging into a system.
[1268] "Conversation Partner" refers to a virtual or real entity that a user selects within the system with whom to interact.
[1269] "Real-time" means that user input is processed almost instantly and responses are returned immediately.
[1270] The present invention is a system that collects, cleans, and stores conversation data in a database. Furthermore, the system generates a large-scale language model for training using the stored conversation data, and provides an appropriate AI model based on a specific conversation partner after the user is authenticated. The system conducts real-time conversations between the user and the AI model, and the AI model generates responses based on the user's input. The system also includes an emotion engine that recognizes the user's emotions and generates responses adapted to those emotions. The emotion engine can recognize emotions from the user's input voice data and facial expression data. The system also includes a means for calculating and distributing rewards to data providers based on the conversation data used.
[1271] Hardware and software used
[1272] The system uses the following hardware and software:
[1273] Server: A server with high performance data storage and computing power (e.g., AWS EC2 instance)
[1274] Database: A relational database (e.g., Amazon RDS, MySQL) to store conversation data.
[1275] Deep learning frameworks for training language models (e.g., TensorFlow, PyTorch)
[1276] Emotion recognition engine: Software for analyzing voice and facial expression data (e.g., OpenCV, emotionAPI)
[1277] User interface: Web or mobile application for data upload and conversation simulation
[1278] Collecting and cleaning conversation data
[1279] The server provides an interface for conversation data providers to log in and upload conversation data in the form of audio, text, or video. Providers upload the data, which is temporarily stored on the server. The server then cleans the uploaded conversation data, removing noise and filtering inappropriate content and personal information, and stores the cleaned data in a database.
[1280] Generating large-scale language models
[1281] The server uses the cleaned conversation data to train large-scale language models, generates models tailored to specific conversation partners, and sets up API endpoints to make them accessible. Training is done using deep learning frameworks such as TensorFlow and PyTorch.
[1282] Emotion engine integration
[1283] The server integrates an emotion engine into the AI model. The emotion engine recognizes emotions from user input data (text, voice, facial expressions) and adjusts responses based on the results. The emotion recognition algorithm evaluates the user's emotional state in real time and generates appropriate responses.
[1284] User authentication and conversation initiation
[1285] The user launches the application on their device and enters their authentication information on a login screen. If authentication is successful, an interface appears where they can select a conversation partner and begin the dialogue. The user inputs text or voice, and the data is sent to the server. The server passes the received input to an AI model and emotion engine, generating a response in real time. The response is displayed on the device, and the user can continue the conversation.
[1286] Reward distribution
[1287] The server tracks users' conversation data and calculates rewards for data providers based on it. Rewards are calculated based on usage time and frequency and distributed appropriately. This increases transparency for data providers and strengthens their motivation to provide data.
[1288] Specific examples
[1289] For example, consider a case where a user simulates a conversation with a psychological counselor. The user inputs, "I've been feeling very stressed lately," and the message is sent to the server. The emotion engine analyzes this input and recognizes the user's emotion as "stress." The AI model generates a basic response, "That must be tough," which the emotion engine then refines by saying, "What situations are causing you stress? Tell me about them." This response is ultimately returned to the user.
[1290] Prompt Sentence Examples
[1291] Example prompt: "Tell me about the latest robotics technology."
[1292] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1293] Step 1:
[1294] Conversation data collection
[1295] Input: Audio, text, or video files provided by the conversation data provider
[1296] Server: A conversation data provider logs in to the system and uploads conversation data through the data upload interface. When the provider selects the data and clicks the upload button, the data is sent to the server and temporarily stored.
[1297] Output: Temporarily saved conversation data
[1298] Specific behavior:
[1299] The user clicks the "Upload conversation data" button.
[1300] The selected file is sent to the server and saved in a temporary folder.
[1301] Step 2:
[1302] Data Preprocessing
[1303] Input: Temporarily saved conversation data
[1304] Server: The uploaded conversation data is cleaned. This process includes removing noise and filtering inappropriate content and personal information. The cleaned data is then stored in a database.
[1305] Output: Cleaned conversation data
[1306] Specific behavior:
[1307] The server adds the uploaded data to a processing queue.
[1308] A cleaning algorithm steps through the data in the queue and applies a noise reduction filter.
[1309] Apply a set of personal information filtering rules to automatically remove inappropriate content.
[1310] Store clean data in the database.
[1311] Step 3:
[1312] AI model generation
[1313] Input: Cleaned conversation data
[1314] Server: Trains large-scale language models using cleaned conversation data, generates models tailored to specific conversation partners, and sets up API endpoints to make them accessible.
[1315] Output: A trained large-scale language model
[1316] Specific behavior:
[1317] The server retrieves the cleaned data from the database.
[1318] Train the model using a deep learning framework (e.g., PyTorch, TensorFlow).
[1319] Save the parameters of the model once it has been trained.
[1320] Create an API endpoint and deploy the model.
[1321] Step 4:
[1322] Emotion engine integration
[1323] Input: A trained large-scale language model
[1324] Server: Integrates an emotion engine into the AI model. The emotion engine contains algorithms for recognizing emotions from user input data (text, voice, facial expressions) and adjusting responses based on the results.
[1325] Output: AI model with integrated emotion engine
[1326] Specific behavior:
[1327] Emotion recognition algorithms perform text and voice analysis.
[1328] Assessing the user's emotional state (e.g., happy, sad, angry) in real time.
[1329] The basic responses generated by the AI model are fine-tuned based on the emotion recognition results to generate the optimal response.
[1330] Step 5:
[1331] User authentication and selection
[1332] Input: Authentication information (email address, password)
[1333] On the device: The user launches the application and accesses the login screen. The user enters their email address and password for authentication. If successful, the conversation partner selection screen is displayed.
[1334] Output: List of conversation partners after successful authentication
[1335] Specific behavior:
[1336] The user launches the application and enters their login ID and password.
[1337] The server verifies the authentication information and, if it matches, starts the session.
[1338] After successful login, a list of conversation partners will be displayed on your device.
[1339] Step 6:
[1340] Conversation Simulation
[1341] Input: User text or voice input
[1342] Terminal: The user selects a conversation partner through the system's interface and starts a conversation. When the user inputs text or voice, the data is sent to the server.
[1343] Server: Receives user input and passes it to the appropriate AI model and emotion engine. The AI model generates a basic response, and the emotion engine adjusts the response based on the user's emotion.
[1344] Output: Adjusted AI response
[1345] Specific behavior:
[1346] The user selects "Psychological Counselor C" and clicks the "Start Dialogue" button.
[1347] The user types "I've been feeling really stressed lately" into the input field and submits.
[1348] The server receives this input and sends it to the AI model.
[1349] The model generates the response "That's tough," which the emotion engine refines to "What situations are stressing you out? Tell me about them."
[1350] The response is returned to the terminal and displayed to the user.
[1351] Step 7:
[1352] Reward distribution
[1353] Input: conversation session data (usage time, usage frequency)
[1354] Server: Tracks user conversation records and calculates rewards for data providers based on them. Calculates rewards based on the time and frequency of use and pays providers.
[1355] Output: Calculated reward points and payment processing
[1356] Specific behavior:
[1357] The server records data for each conversation session.
[1358] A reward calculation algorithm calculates points based on the duration and frequency of use per session.
[1359] Rewards are transferred to conversation data providers through a payment processing service.
[1360] (Application example 2)
[1361] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1362] Customer service in traditional brick-and-mortar stores is heavily dependent on the experience and skills of store staff, making it difficult to maintain consistent service quality. Furthermore, in situations where flexible responses based on customer emotions are required, it is difficult for humans alone to recognize and respond 100% accurately. Furthermore, the appropriate management and effective use of collected conversation data is an issue.
[1363] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting conversation data, means for cleaning the collected conversation data and storing it in a database, means for generating a large-scale language model to be trained using the stored conversation data, means for providing an appropriate AI model based on a specific conversation partner selected by an authenticated user, means for conducting a real-time conversation between the user and the AI model and for the AI model to generate a response based on the user's input, means for generating a response based on the user's emotions using an emotion recognition engine, means for analyzing collected audio and video data using smart glasses and providing appropriate responses to store clerks in real time, and means for calculating and distributing rewards to data providers based on the used conversation data. This enables real-time responses that take customer emotions into consideration and uniform service quality.
[1364] "Conversation Data" means records of audio, text, and video communications between users or between users and the system.
[1365] "Cleaning" is the process of removing noise and organizing data while removing inappropriate content and personal information.
[1366] A "database" is a collection of information that stores collected data in an organized manner and makes it easily searchable and accessible.
[1367] A "large-scale language model" is an artificial intelligence model trained using vast amounts of conversational data to generate natural-sounding language responses based on user input.
[1368] An "AI model" is an artificial intelligence algorithm and its implementation designed to address a specific task or application.
[1369] An "emotion recognition engine" analyzes emotions from a user's voice and video data and adjusts responses based on those emotions.
[1370] "Smart glasses" are a type of wearable device used to acquire and display visual and audio information, and have the ability to present analysis results to users in real time.
[1371] "Reward" means compensation distributed to data providers based on the use of the conversation data they provide.
[1372] "Data providers" are individuals or organizations that upload conversation data to the system.
[1373] The present invention is a system that aims to improve the quality and efficiency of customer support in brick-and-mortar stores, and includes means for collecting and cleaning conversation data, generating large-scale language models, recognizing emotions, and distributing rewards. The following describes in detail the embodiments of the present invention.
[1374] System Program
[1375] 1. Collecting conversation data
[1376] The server collects conversation data through smart glasses worn by store staff. The smart glasses are equipped with a microphone to collect voice data and a camera to capture facial expression data. The collected data is both audio and video.
[1377] 2. Data Preprocessing
[1378] The server receives the raw conversation data sent by the smart glasses and cleans it. The cleaning process includes removing noise and filtering inappropriate content and personal information, improving the quality of the data before storing it in a database.
[1379] 3. Generating large-scale language models
[1380] The server uses the cleaned conversation data to train a large-scale language model, which is trained specifically for specific conversational partners (in this case, the store clerk and the customer) and made accessible through an API, using Python and AI frameworks such as TensorFlow.
[1381] 4. Emotion Recognition Integration
[1382] The server integrates a large-scale language model with an emotion recognition engine, which uses OpenCV and TensorFlow to analyze user emotions from audio and video data. The analysis results are reflected in the generated response.
[1383] 5. User Authentication and Response Generation
[1384] The terminal (smart glasses) performs personal authentication when the store clerk logs in. After logging in, when the store clerk interacts with the customer, voice input and video data are sent to the server. The server receives this data and passes it to an AI model and emotion recognition engine to generate an appropriate response.
[1385] 6. Real-time response
[1386] The server sends the generated response to the smart glasses in real time and displays it to the store clerk. This allows the store clerk to immediately provide an appropriate response to the customer. For example, if a customer asks, "I've been feeling stressed lately," the smart glasses will respond, "That's tough. What situations are making you feel stressed?"
[1387] 7. Reward Distribution
[1388] The server calculates and distributes rewards to data providers (store clerks and stores) based on the collected conversation data. Rewards are determined based on the frequency and quality of data usage.
[1389] Hardware and software used
[1390] Hardware: Smart glasses (e.g., Google Glass Enterprise Edition 2)
[1391] Server software: AWS Lambda, S3, EC2
[1392] Preprocessing and training: Python scripts, TensorFlow
[1393] Emotion recognition: OpenCV, TensorFlow
[1394] Authentication system: OAuth 2.0, Firebase Authentication
[1395] Examples of prompt statements
[1396] For example, if a customer asks, "I've been feeling stressed lately, especially at work...what should I do?" the system might respond with:
[1397] "That's tough. What situations are stressing you out?"
[1398] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1399] Step 1:
[1400] A user (a store clerk) wears smart glasses and serves customers in a store. The smart glasses collect audio and video data in real time and send it to a server. The input is conversation data (audio and video) between the customer and the store clerk, which is then sent to the server.
[1401] Step 2:
[1402] The server receives the transmitted conversation data and performs cleaning processes such as noise removal, filtering of inappropriate content and personal information. The data is then cleaned and stored in a database. The input is raw conversation data, and the output is cleaned data.
[1403] Step 3:
[1404] The server uses the cleaned conversation data to train a large-scale language model. It uses an AI framework such as TensorFlow to generate a model specialized for a specific conversation partner. The input is the cleaned conversation data, and the output is the trained large-scale language model.
[1405] Step 4:
[1406] The server integrates an emotion recognition engine with the trained large-scale language model. This engine uses OpenCV and TensorFlow to recognize user emotions from audio and video data. The input is audio and video data, and the output is recognized emotional information and response adjustments based on it.
[1407] Step 5:
[1408] The smart glasses at the terminal authenticate the store clerk by logging in to the system. During this process, the user's authentication information is sent to the server and verified. The input is the store clerk's authentication information, and the output is the authentication success or failure status.
[1409] Step 6:
[1410] The server receives real-time audio and video data when an authenticated store associate interacts with a customer and passes it to an AI model and emotion recognition engine. The AI model and emotion recognition engine then generate an appropriate response based on the user's input data. The input is real-time conversation data with the customer, and the output is the generated appropriate response.
[1411] Step 7:
[1412] The server sends the generated response to the smart glasses in real time and presents it to the store clerk. The store clerk responds to the customer based on the response. The input is the generated response, and the output is the response displayed on the smart glasses.
[1413] Step 8:
[1414] The server tracks all conversation sessions and calculates and distributes rewards to data providers based on the conversation data used. The input is the conversation session data, and the output is the reward calculation and distribution results.
[1415] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1416] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1417] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1418] [Fourth embodiment]
[1419] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1420] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1421] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1422] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1423] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1424] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1425] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1426] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1427] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1428] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1429] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1430] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1431] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1432] The present invention provides a system that collects, cleans, and stores conversation data in a database, generates a large-scale language model for training using the stored conversation data, and provides an appropriate AI model based on a specific conversation partner after a user is authenticated. The system also conducts real-time conversations between the user and the AI model, and the AI model generates responses based on the user's input. The system also includes a means for calculating and distributing rewards to data providers based on the conversation data used.
[1433] Program processing
[1434] 1. Collecting conversation data
[1435] Server: Provides an interface for conversation data providers (celebrities and professionals) to upload their conversation data after logging in. When providers upload audio, text, or video data, the data is temporarily stored.
[1436] 2. Data Preprocessing
[1437] Server: Analyzes the received conversation data, removes noise and unnecessary parts, and filters out personal information. The preprocessed data is saved in the database as cleaned data.
[1438] 3. AI model generation
[1439] Server: Uses the cleaned data to train large-scale language models, which are then tuned to specific conversational partners and set up API endpoints to make them accessible.
[1440] 4. User Authentication and Selection
[1441] Terminal: Provides an authentication interface for users to access the system through the application and log in. Once authenticated, users can select a specific conversation partner from a list.
[1442] 5. Conversation Simulation
[1443] Terminal: The terminal displays an input form on the interface to start a conversation with a conversation partner of the user's choice. The user enters the information and sends it to the server.
[1444] Server: Receives user input and passes it to the appropriate AI model, which generates a response and sends it back to the device.
[1445] Terminal: Displays responses from the server and provides an environment in which the user can continue the dialogue.
[1446] 6. Reward Distribution
[1447] Server: Tracks the user's conversation time and frequency of use, and calculates the reward for the conversation data provider. Once the reward is calculated, payment is processed based on the provider's account information.
[1448] Specific examples
[1449] For example, consider the case where a famous actor B provides audio data about his or her stage experience. This data is collected, cleaned to remove personal information, and then stored in a database. Next, a large-scale language model is trained based on Actor B's audio data to generate an AI model that can converse with Actor B.
[1450] A user logs into this system and selects "Simulated dialogue with actor B." When the user inputs "What was your most recent stage challenge?" to actor B, the input is sent to the server, and the AI model generates a response such as "My most recent stage challenge was trying out a new role," and sends it back to the user.
[1451] When the conversation ends, the server calculates the reward for Actor B based on the time and frequency of use and transfers it to the system. In this way, users can advance their language learning through a realistic dialogue experience, and the provider of the conversation data can also earn rewards.
[1452] This system not only allows language learners to stay motivated and improve their skills effectively, but also allows conversation data providers to earn revenue in new ways.
[1453] The processing flow will be explained below.
[1454] Understood. Below is the process flow broken down into specific steps.
[1455] Step 1:
[1456] Server: Provides an authentication interface for conversation data providers (e.g., famous actor B) to log in. The conversation data provider logs in by entering the correct authentication information.
[1457] Step 2:
[1458] Server: After successful authentication, the server displays a conversation data upload interface to the conversation data provider, where the provider can select and upload conversation data in audio, text, or video format.
[1459] Step 3:
[1460] Server: Temporarily stores the uploaded data, checks the format and content of the data, and if the format is correct, sends the data to the cleaning process.
[1461] Step 4:
[1462] Server: The cleaning process removes noise and unnecessary parts from the conversation data and filters out personal information, such as the patient's name and contact details.
[1463] Step 5:
[1464] Server: Organizes the cleaned data and stores it in a database. When storing the data, it categorizes it by conversation topic, person, date and time, etc.
[1465] Step 6:
[1466] Server: Uses the cleaned conversation data to train a large-scale language model (LLM). For example, it trains the model based on actor B's data and builds a response generation model specialized for that person.
[1467] Step 7:
[1468] Server: Evaluates the trained model on the test data, adjusts the model parameters if necessary, saves the completed model, and sets up the appropriate API endpoints.
[1469] Step 8:
[1470] Terminal: The user launches the application and accesses the login screen. The user enters their credentials and logs in.
[1471] Step 9:
[1472] Server: Checks the user's authentication information, and if authentication is successful, displays the user dashboard, where the user can select the conversation partner they want to simulate (e.g., Actor B).
[1473] Step 10:
[1474] Terminal: Displays an interface for the user to select a simulated conversation partner. The user makes the selection and sends the selection information to the server.
[1475] Step 11:
[1476] Server: Based on the user's selection, prepares the corresponding AI model (e.g., actor B's model) and sends interface information to the terminal to start the conversation.
[1477] Step 12:
[1478] Terminal: Sends a text message typed by the user to the server. For example, "What's your latest stage challenge?"
[1479] Step 13:
[1480] Server: Receives user input and passes it to the appropriate AI model, which then generates a response, for example, "My most recent stage challenge was taking on a new role."
[1481] Step 14:
[1482] Server: Sends the generated response to the user's device.
[1483] Step 15:
[1484] Terminal: Receives the response from the server and displays it in the user interface. If the user wants to continue the interaction, they can input again.
[1485] Step 16:
[1486] Server: After the user's simulation is completed, the server calculates the reward for the conversation data provider based on the simulation time and number of times used.
[1487] Step 17:
[1488] Server: Carries out the payment procedures to distribute the calculated reward to the conversation data provider. The reward is transferred based on the provider's account information.
[1489] The above is the specific processing flow of this system.
[1490] Example 1
[1491] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1492] Conventional conversation data collection and processing systems have insufficient quality control of collected data, particularly due to the inclusion of noise and personal information, which limits its use. Furthermore, when training AI models using conversation data, it is difficult to adjust them to specific conversation partners, making it difficult to provide real-time responses to users. Furthermore, reward distribution to data providers is sometimes not transparent and efficient, resulting in a lack of incentives for providers.
[1493] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1494] In this invention, the server includes a means for collecting conversation data, a means for cleaning the collected conversation data and storing it in a database, and a means for generating a model to be trained using the stored conversation data. This enables appropriate preprocessing of the data and the generation of an AI model tailored to a specific conversation partner using high-quality training data. The server also includes a means for selecting a specific conversation partner after the user is authenticated, a means for conducting a conversation between the user and the model and for the model to generate a response based on the user's input, and a means for calculating and distributing rewards to data providers based on the used conversation data. This allows users to receive appropriate responses in real time and distributes rewards fairly and quickly to data providers, improving the efficiency and satisfaction of the entire system.
[1495] "Conversation data" refers to audio data, text data, or video data provided to the system by users or conversation data providers.
[1496] "Cleaning" is the process of removing noise, inappropriate content, personal information, etc. from collected conversation data to improve data quality before storing it in a database.
[1497] A "database" is a storage within the system that efficiently and securely stores cleaned conversation data for later use as training data.
[1498] "Training" is the process of using cleaned conversational data to train an AI model and improve its pattern recognition capabilities and prediction accuracy.
[1499] A "model" is a collection of AI structures and algorithms trained on conversational data that are used to generate appropriate responses to user input.
[1500] A "specific conversation partner" is a virtual conversation partner for a dialogue that is selected by the user based on various categories and profiles provided on the system.
[1501] "Real-time" means that there is a very short delay between when a user makes an input and when the AI model returns a response, and refers to a situation in which the interaction takes place immediately.
[1502] A "response" is a response, such as text or voice, that an AI model generates based on user input.
[1503] "Remuneration" is the compensation calculated and paid to the conversation data provider based on the degree and frequency of use of the provided data.
[1504] "Distribution" is the process of appropriately paying the calculated reward based on the registered information of the conversation data provider.
[1505] "Interface" refers to the screens, input forms, and operating means used by conversation data providers when uploading data or by users when accessing and operating the system.
[1506] To implement this invention, the server, terminal, and user each play specific roles. The main hardware used includes a database server, web server, and user terminal (PC or smartphone), while the software required includes a framework for training large-scale language models (e.g., TensorFlow, PyTorch), database software (e.g., MySQL, PostgreSQL), and a web service API.
[1507] The server provides a means for collecting conversation data. Specifically, it provides a user interface for conversation data providers to log in and upload audio data, text data, and video data. The device supports a means for providers to upload data, allowing them to select and send data in various formats (e.g., .mp3, .txt, .mp4). Uploaded data is temporarily stored on the server.
[1508] The server analyzes the collected conversation data, removes noise and unnecessary parts, and filters out personal information to clean it. The cleaned data is then stored in a database, which later serves as the foundation for training AI models.
[1509] The server then uses the cleaned conversation data to train a large-scale language model using a large-scale language model framework, processing the data in batches to learn the model weights, resulting in an AI model that corresponds to a specific conversation partner. The model is then made accessible through an API endpoint.
[1510] To access the system, a user uses a terminal to enter a username and password into a login interface. After successful authentication, the server provides the user with a list of specific available conversation partners from which the user can select a conversation partner.
[1511] The device then displays an interface for the user to converse with the selected conversation partner. The user enters a question or message into the input form and clicks the send button. For example, a question such as "What was your most recent stage challenge?" is sent to the server.
[1512] The server receives user input, passes it to the trained AI model, and generates a response, which is then sent back to the device, where it is displayed to the user, allowing the user to continuously engage in real-time dialogue with the AI model.
[1513] Finally, the server tracks the duration and frequency of users' conversation data usage and calculates rewards for conversation data providers based on that data. The calculated rewards are distributed based on the providers' registered account information.
[1514] As a concrete example, consider a case where a user selects a dialogue simulation with actor B and asks, "What was the most moving moment on stage?" This question is transmitted to the AI model via the server, and the AI model generates a response such as, "The most moving moment on stage was when the audience shed tears," and sends it back to the user.
[1515] This system allows users to advance their language learning through realistic conversational experiences, and enables conversation data providers to earn revenue in new ways.
[1516] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1517] Step 1:
[1518] The server verifies the login of the conversation data provider and, after successful authentication, provides an interface for data upload.
[1519] Input: Provider username and password.
[1520] Specific operation: Query the database to verify the provider's authentication information, and if authentication is successful, start a login session.
[1521] Output: Login session token and data upload interface.
[1522] Step 2:
[1523] Using the terminal, the conversation data provider selects audio data, text data, or video data and clicks the upload button.
[1524] Input: Data files selected by the provider (e.g. .mp3, .txt, .mp4).
[1525] Specific operation: A file selection dialog is displayed, and after the provider selects a file, they press the upload button.
[1526] Output: The selected data file is sent to the server.
[1527] Step 3:
[1528] The server stores the uploaded data in temporary storage and checks the integrity of the data.
[1529] Input: Uploaded data file.
[1530] What it does: Saves the file to temporary storage and verifies the data integrity by checking the file format and the first few bytes of the data.
[1531] Output: Data file with integrity checked.
[1532] Step 4:
[1533] The server analyzes the received conversation data, removes noise and unnecessary parts, and filters out personal information to clean it.
[1534] Input: Data file whose integrity has been checked.
[1535] Specific operation: For audio data, a noise reduction algorithm is applied, and for text data, filtering is performed to mask specific keywords.
[1536] Output: Cleaned data.
[1537] Step 5:
[1538] The server organizes and stores the cleaned data in a database.
[1539] Input: Cleaned data.
[1540] What it does: Inserts data and creates indexes according to the appropriate schema in the database.
[1541] Output: Cleaned data stored in a database.
[1542] Step 6:
[1543] The server uses the cleaned conversational data to train a large-scale language model.
[1544] Input: Cleaned data in the database.
[1545] Specific behavior: Use a large-scale language model framework (e.g., TensorFlow, PyTorch) to process data in batches and update the model weights.
[1546] Output: A trained AI model.
[1547] Step 7:
[1548] The server tunes the trained model to correspond to specific conversation partners and sets up an API endpoint to make it accessible externally.
[1549] Input: A trained AI model.
[1550] Specific actions: Adjust model parameters, configure API endpoints, and deploy.
[1551] Output: An accessible API endpoint for the AI model.
[1552] Step 8:
[1553] The device presents the user with a login interface, prompting them to enter their username and password, and upon successful authentication, providing an interface for selecting a specific conversation partner.
[1554] Input: Username and Password.
[1555] Specific behavior: The user enters information into the login form and clicks the login button.
[1556] Output: Conversation partner selection interface for authenticated user.
[1557] Step 9:
[1558] The terminal displays a screen for starting a conversation based on the conversation partner selected by the user.
[1559] Input: User's selected conversation partner information.
[1560] Specific behavior: Displays a conversation interface and provides an input form.
[1561] Output: A form for the user to enter.
[1562] Step 10:
[1563] The user enters a question or message into the input form and clicks the send button.
[1564] Input: User question or message (e.g., "What was your most recent stage challenge?").
[1565] Specific behavior: The user enters text and clicks the submit button.
[1566] Output: The entered question or message is sent to the server.
[1567] Step 11:
[1568] The server receives user input and passes it to a trained AI model to generate a response.
[1569] Input: The user's typed question or message.
[1570] What it does: Tokenizes the user's message and passes it to the AI model for processing.
[1571] Output: The generated response of the AI model.
[1572] Step 12:
[1573] The terminal displays responses from the server to the user, providing an interface that allows the conversation to continue.
[1574] Input: The response sent back from the server.
[1575] Specific behavior: Displays the response text and persists the interface so the user can enter it again.
[1576] Output: The response displayed to the user and an input form for further interaction.
[1577] Step 13:
[1578] The server tracks the time and frequency of users' conversation data usage and calculates rewards for conversation data providers based on that data.
[1579] Input: User conversation time, usage frequency data.
[1580] Specific operation: Analyzes conversation logs and calculates rewards based on an algorithm.
[1581] Output: Calculated reward data.
[1582] Step 14:
[1583] The server distributes the calculated reward based on the provider's registered account information.
[1584] Input: Remuneration data and provider account information.
[1585] Specific operation: The reward is transferred to the provider's account.
[1586] Output: Rewards deposited into the provider's account.
[1587] The above are the specific processing steps of the program in this system.
[1588] (Application example 1)
[1589] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1590] The present invention relates to a technology and system that provides dialogue simulations to meet the needs of users who want to have real-time dialogues with specific conversation partners. It also aims to solve the problem of simultaneously providing a reward distribution mechanism for dialogue data providers. This allows users to enhance their learning by engaging in dialogues that include specialized knowledge, and data providers to gain new revenue sources.
[1591] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1592] In this invention, the server includes means for collecting conversation data, means for cleaning the collected conversation data and storing it in a database, means for generating a large-scale language model for training using the stored conversation data, means for providing an appropriate AI model based on a specific conversation partner selected by an authenticated user, means for conducting a conversation between the user and the AI model in real time and for the AI model to generate a response based on the user's input, means for calculating and distributing rewards to data providers based on the used conversation data, means for providing an interface for the user to start a dialogue simulation, and means for tracking the user's conversation data and transferring rewards to dialogue data providers through the system. This allows users to simulate a dialogue with a specific conversation partner in real time and effectively learn and exchange information, while also allowing dialogue data providers to receive fair rewards.
[1593] "Conversation Data" means information in the form of audio, text, or video exchanged between Users or between Users and the System.
[1594] "Cleaning" refers to the process of removing noise, inappropriate content, and personal information from collected conversation data.
[1595] "Database" refers to a system that accumulates and manages cleaned conversation data.
[1596] A "large-scale language model" refers to a natural language processing model trained on a massive amount of conversational data.
[1597] "AI model" refers to a large-scale language model tuned based on a specific conversation partner.
[1598] "Real-time" refers to a situation where there is almost no delay and processing or response is immediate.
[1599] "Conversation Partner" refers to a specific expert, celebrity, or AI model thereof with whom the user wishes to have a conversation.
[1600] "Interface" refers to the screen and input means that users use to operate the system.
[1601] "Tracking" refers to the act of following and recording user behavior and conversation data.
[1602] "Reward" refers to money or other consideration paid by the system to the dialogue data provider.
[1603] The present invention is a system that allows users to have real-time conversations with specific conversation partners. The system provides a series of functions including collection, cleaning, and storage of conversation data, generation of AI models, user authentication, real-time conversation execution, and reward distribution.
[1604] System program generation
[1605] Data collection and cleaning
[1606] The server provides an interface for conversation data providers (such as experts and celebrities) to upload audio, text, and video data after they log in. This data is temporarily stored and analyzed to filter out noise, unnecessary parts, and personal information. Once preprocessing is complete, the data is stored in a database as cleaned data.
[1607] Generating AI models
[1608] The server uses the cleaned data to train large-scale language models, which are then tuned to specific conversational partners and set up API endpoints to make them accessible.
[1609] User authentication and conversation simulation
[1610] The terminal (smartphone) provides an authentication interface for users to access and log in to the system. Once authentication is complete, the user is presented with an interface to select a specific conversation partner and begin the dialogue simulation. The user's input is sent to the server, and the corresponding AI model generates a response and sends it back to the terminal. The server performs this process in real time, providing an environment in which the user can continue the dialogue.
[1611] Reward distribution
[1612] The server tracks the user's conversation time and frequency of use, calculates the reward for the conversation data provider, and then processes the payment based on the provider's account information.
[1613] Hardware and software used
[1614] Hardware: A server with a high-performance CPU and GPU, and a smartphone (iOS or Android device)
[1615] Software: Applications running on edge devices (Swift or Kotlin), Python scripts for data analysis and filtering, database management system (MySQL), large-scale language model (OpenAI GPT), API server (Flask or Django)
[1616] Specific examples
[1617] For example, if a user logs in to the system and selects a simulated conversation with a specific expert, the following process will occur: When the user types, "Tell me about your new business idea," the server receives this input and passes it to an AI model trained as a specific expert. The AI model generates a response: "The key to a new business idea is market research and finding a niche." The generated response is sent back to the terminal, and the user can use it to ask further questions.
[1618] Prompt Sentence Examples
[1619] "If a user asks, 'Tell me about a new business idea,' generate a response from an AI trained as a business expert."
[1620] This allows users to enhance their learning by engaging in conversations that include specialized knowledge, and provides conversation data providers with a new source of revenue.
[1621] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1622] Program processing steps
[1623] Step 1:
[1624] Conversation data collection
[1625] Server: When a conversation data provider logs in, it provides an interface for uploading voice data, text data, and video data. Once the provider uploads the data, it is temporarily stored.
[1626] Input: Conversation data (audio, text, video)
[1627] Output: Temporarily saved conversation data
[1628] Specific operation: After authenticating the provider, the server stores the uploaded data in a specific directory on the server.
[1629] Step 2:
[1630] Data Preprocessing
[1631] Server: Analyzes the collected conversation data, filters out noise, unnecessary parts, and personal information, and stores the cleaned data in a database.
[1632] Input: Temporarily saved conversation data
[1633] Output: Cleaned conversation data
[1634] What it does: It uses a Python script to denoise the audio data, sanitize the text, and filter out personal information, then stores it in a MySQL database.
[1635] Step 3:
[1636] Generating AI models
[1637] Server: Trains large-scale language models using cleaned data, tunes AI models based on specific conversational partners, and sets up API endpoints for access.
[1638] Input: Cleaned conversation data
[1639] Output: A tuned AI model
[1640] What it does: Train and tune large-scale language models using Python and TensorFlow. Set up API endpoints using Flask or Django.
[1641] Step 4:
[1642] User authentication and selection
[1643] Terminal (smartphone): Provides an authentication interface for users to access the system through the app and log in. Once authenticated, users can select a specific conversation partner.
[1644] Input: User login information, conversation partner selection
[1645] Output: Preparation for starting the conversation simulation
[1646] Specific operation: The login screen is displayed on the device using Swift or Kotlin, and the user's input information is sent to the server for authentication. After authentication, a list of conversation partners is displayed.
[1647] Step 5:
[1648] Conversation Simulation
[1649] Terminal (smartphone): Displays an interface that allows the user to start a conversation with a selected conversation partner and sends the user's input to the server.
[1650] Server: Passes user input to the AI model, generates a response, and sends it back to the device, which displays the response to the user.
[1651] Input: User's spoken input
[1652] Output: Response by the AI model
[1653] Specific operation: Receives user input and sends it to the server, which passes the input to the AI model and returns the generated response to the mobile device, which displays the response.
[1654] Step 6:
[1655] Reward distribution
[1656] Server: Tracks user conversation time and frequency of use, calculates rewards for conversation data providers, and processes payments.
[1657] Input: Conversation transcript and usage data
[1658] Output: Reward calculation and payment
[1659] Specific operation: A Python script aggregates conversation time and frequency of use, calculates the reward for the provider, and processes the payment to the provider's registered account.
[1660] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1661] The present invention provides a system that collects, cleans, and stores conversation data in a database. It also generates a large-scale language model for training using the stored conversation data, and provides an appropriate AI model based on a specific conversation partner after a user is authenticated. It also conducts real-time conversations between the user and the AI model, and the AI model generates responses based on the user's input. It also includes an emotion engine that recognizes the user's emotions and generates responses adapted to those emotions. This emotion engine can recognize emotions from the user's input voice data and facial expression data. It also includes a means for calculating and distributing rewards to data providers based on the conversation data used.
[1662] Program processing
[1663] 1. Collecting conversation data
[1664] Server: Conversation data providers log in to the system and upload data in the form of audio, text, or video through the conversation data upload interface. This data is temporarily stored on the server.
[1665] 2. Data Preprocessing
[1666] Server: The uploaded conversation data is cleaned. This process includes removing noise and filtering inappropriate content and personal information. The cleaned data is then stored in a database.
[1667] 3. AI model generation
[1668] Server: Trains large-scale language models using cleaned conversation data, trains models specific to specific conversation partners, and sets up API endpoints for access.
[1669] 4. Emotion engine integration
[1670] Server: Integrates an emotion engine into the AI model. The emotion engine contains algorithms for recognizing emotions from user input (text, voice, facial expressions, etc.) and adjusting responses based on the recognized emotions.
[1671] 5. User Authentication and Selection
[1672] Terminal: The user launches the application and accesses the login screen. The user enters their credentials and logs in.
[1673] 6. Conversation Simulation
[1674] Terminal: The user selects a conversation partner through the system's interface and starts a dialogue. The terminal transmits the text and voice data entered by the user to the server.
[1675] Server: Receives user input and passes it to the appropriate AI model and emotion engine. The AI model generates a basic response, and the emotion engine tailors the response based on the user's emotion. For example, if a user makes an input that sounds like they're in distress, the emotion engine will recognize that and generate a response like, "You seem distressed. Is there anything I can help you with?"
[1676] Terminal: displays the adjusted response to the user and allows the user to continue typing.
[1677] 7. Reward Distribution
[1678] Server: Tracks user conversation session data and calculates rewards for data providers based on the time and frequency of use, and pays providers.
[1679] Specific examples
[1680] For example, consider the case where psychological counselor C provides audio data from a session. Counselor C's data is collected, cleaned, and stored in a database. Next, a large-scale language model is trained based on psychological counselor C's data to build an AI model that generates specific responses.
[1681] A user logs into the system and selects "Simulated conversation with psychological counselor C." When the user enters "I've been feeling very stressed lately," the message is sent to the server. The emotion engine analyzes this input and recognizes the user's emotions as "stress" or "anxiety." The AI model generates a basic response of "That must be tough," which the emotion engine refines to "What situations are causing you stress? Tell me about them." This response is ultimately returned to the user.
[1682] This system allows users to have a realistic conversational experience, and allows them to receive psychological support while learning a language. Conversation data providers are also rewarded based on the use of their data.
[1683] The processing flow will be explained below.
[1684] Understood. Below is the process flow broken down into specific steps.
[1685] Step 1:
[1686] Server: Provides an authentication interface for conversation data providers to log in. The conversation data provider logs in by entering the correct authentication information.
[1687] Step 2:
[1688] Server: After successful authentication, the server displays an interface for uploading conversation data to the data provider, allowing the provider to select and upload data in audio, text, or video format.
[1689] Step 3:
[1690] Server: Temporarily stores the uploaded data, checks the format and content of the data, and if the format is correct, sends the data through the cleaning process.
[1691] Step 4:
[1692] Server: The cleaning process removes noise and unnecessary parts from the conversation data and filters out personal information, such as the patient's name and contact details.
[1693] Step 5:
[1694] Server: Organizes the cleaned data and stores it in a database. When storing the data, it categorizes it by conversation topic, person, date and time, etc.
[1695] Step 6:
[1696] Server: Trains a large-scale language model (LLM) using the cleaned conversation data. For example, it trains the model based on the data of psychological counselor C and builds a response generation model specialized for that person.
[1697] Step 7:
[1698] Server: Evaluates the trained model on the test data, adjusts the model parameters if necessary, saves the completed model, and sets up the appropriate API endpoints.
[1699] Step 8:
[1700] Server: Integrates the emotion engine into the AI model. The emotion engine contains algorithms to recognize emotions from user input (text, voice, facial expressions, etc.) and make appropriate adjustments.
[1701] Step 9:
[1702] Terminal: The user launches the application and accesses the login screen. The user enters their credentials and logs in.
[1703] Step 10:
[1704] Server: Checks the user's authentication information, and if authentication is successful, displays the user dashboard, where the user can select the conversation partner they want to simulate (e.g., psychological counselor C).
[1705] Step 11:
[1706] Terminal: Provides an interface for the user to select a simulated conversation partner. The user makes the selection and sends the selection information to the server.
[1707] Step 12:
[1708] Server: To start a simulation based on the user's selection, prepare the corresponding AI model (psychological counselor C's model) and send information to the user's device to optimize performance.
[1709] Step 13:
[1710] Device: Sends text or voice data entered by the user to the server. For example, you might enter, "I've been feeling really stressed lately."
[1711] Step 14:
[1712] Server: Receives user input and passes it to the appropriate AI model. The emotion engine also analyzes this input. The AI model generates a basic response such as "That must be tough," while the emotion engine recognizes emotions such as "stress" and "anxiety" and adjusts the response to "What situations are stressing you out? Tell me about them."
[1713] Step 15:
[1714] Server: Sends the generated response to the user's device.
[1715] Step 16:
[1716] Terminal: Receives the response from the server and displays it in the user interface. The user can then input again if they wish to continue the interaction.
[1717] Step 17:
[1718] Server: After the user's simulation is completed, the server calculates the reward for the data provider based on the simulation time and number of times used, and pays the provider.
[1719] The above is a specific processing flow of the embodiment of this system.
[1720] Example 2
[1721] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1722] Conventional dialogue systems have the following problems. First, the quality of the collected conversation data is inconsistent, resulting in low reliability of the data used for training. Second, they lack the ability to understand the user's emotions and generate responses accordingly, limiting the user experience. Furthermore, the distribution of rewards to data providers is unclear, resulting in a lack of incentive to provide data. These problems must be resolved.
[1723] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1724] In this invention, the server includes means for collecting conversation data, means for cleaning the collected conversation data and storing it in a database, means for generating a large-scale language model trained using the stored conversation data, means for integrating an emotion engine that recognizes user emotions and adjusts responses, and means for calculating and distributing rewards to data providers based on the conversation data used. This makes it possible to generate a reliable large-scale language model using high-quality data and provide appropriate responses according to the user's emotions. Furthermore, transparency of rewards for conversation data providers can be ensured, improving motivation for data provision.
[1725] "Conversational Data" refers to the content of conversations between humans recorded in audio, text, or video format.
[1726] "Means of collection" refers to the mechanism by which user-submitted conversation data is incorporated into the system.
[1727] "Cleaning methods" refers to the process used to remove noise, inappropriate content, and personal information from collected conversation data.
[1728] "Means for storing in a database" refers to the technology used to store the cleaned conversation data as part of a database.
[1729] "Means for generating large-scale language models" refers to a method for using cleaned conversational data to create trained language models based on natural language processing.
[1730] "Means of integrating an emotion engine" refers to the process of incorporating a system into an AI model to recognize emotions from user input data and adjust the model's response based on those emotions.
[1731] "Means for calculating and distributing rewards" refers to a system that calculates and appropriately distributes rewards to data providers based on the amount and frequency of conversation data used.
[1732] "User authentication" refers to the process of verifying a user's identity when logging into a system.
[1733] "Conversation Partner" refers to a virtual or real entity that a user selects within the system with whom to interact.
[1734] "Real-time" means that user input is processed almost instantly and responses are returned immediately.
[1735] The present invention is a system that collects, cleans, and stores conversation data in a database. Furthermore, the system generates a large-scale language model for training using the stored conversation data, and provides an appropriate AI model based on a specific conversation partner after the user is authenticated. The system conducts real-time conversations between the user and the AI model, and the AI model generates responses based on the user's input. The system also includes an emotion engine that recognizes the user's emotions and generates responses adapted to those emotions. The emotion engine can recognize emotions from the user's input voice data and facial expression data. The system also includes a means for calculating and distributing rewards to data providers based on the conversation data used.
[1736] Hardware and software used
[1737] The system uses the following hardware and software:
[1738] Server: A server with high performance data storage and computing power (e.g., AWS EC2 instance)
[1739] Database: A relational database (e.g., Amazon RDS, MySQL) to store conversation data.
[1740] Deep learning frameworks for training language models (e.g., TensorFlow, PyTorch)
[1741] Emotion recognition engine: Software for analyzing voice and facial expression data (e.g., OpenCV, emotionAPI)
[1742] User interface: Web or mobile application for data upload and conversation simulation
[1743] Collecting and cleaning conversation data
[1744] The server provides an interface for conversation data providers to log in and upload conversation data in the form of audio, text, or video. Providers upload the data, which is temporarily stored on the server. The server then cleans the uploaded conversation data, removing noise and filtering inappropriate content and personal information, and stores the cleaned data in a database.
[1745] Generating large-scale language models
[1746] The server uses the cleaned conversation data to train large-scale language models, generates models tailored to specific conversation partners, and sets up API endpoints to make them accessible. Training is done using deep learning frameworks such as TensorFlow and PyTorch.
[1747] Emotion engine integration
[1748] The server integrates an emotion engine into the AI model. The emotion engine recognizes emotions from user input data (text, voice, facial expressions) and adjusts responses based on the results. The emotion recognition algorithm evaluates the user's emotional state in real time and generates appropriate responses.
[1749] User authentication and conversation initiation
[1750] The user launches the application on their device and enters their authentication information on a login screen. If authentication is successful, an interface appears where they can select a conversation partner and begin the dialogue. The user inputs text or voice, and the data is sent to the server. The server passes the received input to an AI model and emotion engine, generating a response in real time. The response is displayed on the device, and the user can continue the conversation.
[1751] Reward distribution
[1752] The server tracks users' conversation data and calculates rewards for data providers based on it. Rewards are calculated based on usage time and frequency and distributed appropriately. This increases transparency for data providers and strengthens their motivation to provide data.
[1753] Specific examples
[1754] For example, consider a case where a user simulates a conversation with a psychological counselor. The user inputs, "I've been feeling very stressed lately," and the message is sent to the server. The emotion engine analyzes this input and recognizes the user's emotion as "stress." The AI model generates a basic response, "That must be tough," which the emotion engine then refines by saying, "What situations are causing you stress? Tell me about them." This response is ultimately returned to the user.
[1755] Prompt Sentence Examples
[1756] Example prompt: "Tell me about the latest robotics technology."
[1757] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1758] Step 1:
[1759] Conversation data collection
[1760] Input: Audio, text, or video files provided by the conversation data provider
[1761] Server: A conversation data provider logs in to the system and uploads conversation data through the data upload interface. When the provider selects the data and clicks the upload button, the data is sent to the server and temporarily stored.
[1762] Output: Temporarily saved conversation data
[1763] Specific behavior:
[1764] The user clicks the "Upload conversation data" button.
[1765] The selected file is sent to the server and saved in a temporary folder.
[1766] Step 2:
[1767] Data Preprocessing
[1768] Input: Temporarily saved conversation data
[1769] Server: The uploaded conversation data is cleaned. This process includes removing noise and filtering inappropriate content and personal information. The cleaned data is then stored in a database.
[1770] Output: Cleaned conversation data
[1771] Specific behavior:
[1772] The server adds the uploaded data to a processing queue.
[1773] A cleaning algorithm steps through the data in the queue and applies a noise reduction filter.
[1774] Apply a set of personal information filtering rules to automatically remove inappropriate content.
[1775] Store clean data in the database.
[1776] Step 3:
[1777] AI model generation
[1778] Input: Cleaned conversation data
[1779] Server: Trains large-scale language models using cleaned conversation data, generates models tailored to specific conversation partners, and sets up API endpoints to make them accessible.
[1780] Output: A trained large-scale language model
[1781] Specific behavior:
[1782] The server retrieves the cleaned data from the database.
[1783] Train the model using a deep learning framework (e.g., PyTorch, TensorFlow).
[1784] Save the parameters of the model once it has been trained.
[1785] Create an API endpoint and deploy the model.
[1786] Step 4:
[1787] Emotion engine integration
[1788] Input: A trained large-scale language model
[1789] Server: Integrates an emotion engine into the AI model. The emotion engine contains algorithms for recognizing emotions from user input data (text, voice, facial expressions) and adjusting responses based on the results.
[1790] Output: AI model with integrated emotion engine
[1791] Specific behavior:
[1792] Emotion recognition algorithms perform text and voice analysis.
[1793] Assessing the user's emotional state (e.g., happy, sad, angry) in real time.
[1794] The basic responses generated by the AI model are fine-tuned based on the emotion recognition results to generate the optimal response.
[1795] Step 5:
[1796] User authentication and selection
[1797] Input: Authentication information (email address, password)
[1798] On the device: The user launches the application and accesses the login screen. The user enters their email address and password for authentication. If successful, the conversation partner selection screen is displayed.
[1799] Output: List of conversation partners after successful authentication
[1800] Specific behavior:
[1801] The user launches the application and enters their login ID and password.
[1802] The server verifies the authentication information and, if it matches, starts the session.
[1803] After successful login, a list of conversation partners will be displayed on your device.
[1804] Step 6:
[1805] Conversation Simulation
[1806] Input: User text or voice input
[1807] Terminal: The user selects a conversation partner through the system's interface and starts a conversation. When the user inputs text or voice, the data is sent to the server.
[1808] Server: Receives user input and passes it to the appropriate AI model and emotion engine. The AI model generates a basic response, and the emotion engine adjusts the response based on the user's emotion.
[1809] Output: Adjusted AI response
[1810] Specific behavior:
[1811] The user selects "Psychological Counselor C" and clicks the "Start Dialogue" button.
[1812] The user types "I've been feeling really stressed lately" into the input field and submits.
[1813] The server receives this input and sends it to the AI model.
[1814] The model generates the response "That's tough," which the emotion engine refines to "What situations are stressing you out? Tell me about them."
[1815] The response is returned to the terminal and displayed to the user.
[1816] Step 7:
[1817] Reward distribution
[1818] Input: conversation session data (usage time, usage frequency)
[1819] Server: Tracks user conversation records and calculates rewards for data providers based on them. Calculates rewards based on the time and frequency of use and pays providers.
[1820] Output: Calculated reward points and payment processing
[1821] Specific behavior:
[1822] The server records data for each conversation session.
[1823] A reward calculation algorithm calculates points based on the duration and frequency of use per session.
[1824] Rewards are transferred to conversation data providers through a payment processing service.
[1825] (Application example 2)
[1826] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1827] Customer service in traditional brick-and-mortar stores is heavily dependent on the experience and skills of store staff, making it difficult to maintain consistent service quality. Furthermore, in situations where flexible responses based on customer emotions are required, it is difficult for humans alone to recognize and respond 100% accurately. Furthermore, the appropriate management and effective use of collected conversation data is an issue.
[1828] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting conversation data, means for cleaning the collected conversation data and storing it in a database, means for generating a large-scale language model to be trained using the stored conversation data, means for providing an appropriate AI model based on a specific conversation partner selected by an authenticated user, means for conducting a real-time conversation between the user and the AI model and for the AI model to generate a response based on the user's input, means for generating a response based on the user's emotions using an emotion recognition engine, means for analyzing collected audio and video data using smart glasses and providing appropriate responses to store clerks in real time, and means for calculating and distributing rewards to data providers based on the used conversation data. This enables real-time responses that take customer emotions into consideration and uniform service quality.
[1829] "Conversation Data" means records of audio, text, and video communications between users or between users and the system.
[1830] "Cleaning" is the process of removing noise and organizing data while removing inappropriate content and personal information.
[1831] A "database" is a collection of information that stores collected data in an organized manner and makes it easily searchable and accessible.
[1832] A "large-scale language model" is an artificial intelligence model trained using vast amounts of conversational data to generate natural-sounding language responses based on user input.
[1833] An "AI model" is an artificial intelligence algorithm and its implementation designed to address a specific task or application.
[1834] An "emotion recognition engine" analyzes emotions from a user's voice and video data and adjusts responses based on those emotions.
[1835] "Smart glasses" are a type of wearable device used to acquire and display visual and audio information, and have the ability to present analysis results to users in real time.
[1836] "Reward" means compensation distributed to data providers based on the use of the conversation data they provide.
[1837] "Data providers" are individuals or organizations that upload conversation data to the system.
[1838] The present invention is a system that aims to improve the quality and efficiency of customer support in brick-and-mortar stores, and includes means for collecting and cleaning conversation data, generating large-scale language models, recognizing emotions, and distributing rewards. The following describes in detail the embodiments of the present invention.
[1839] System Program
[1840] 1. Collecting conversation data
[1841] The server collects conversation data through smart glasses worn by store staff. The smart glasses are equipped with a microphone to collect voice data and a camera to capture facial expression data. The collected data is both audio and video.
[1842] 2. Data Preprocessing
[1843] The server receives the raw conversation data sent by the smart glasses and cleans it. The cleaning process includes removing noise and filtering inappropriate content and personal information, improving the quality of the data before storing it in a database.
[1844] 3. Generating large-scale language models
[1845] The server uses the cleaned conversation data to train a large-scale language model, which is trained specifically for specific conversational partners (in this case, the store clerk and the customer) and made accessible through an API, using Python and AI frameworks such as TensorFlow.
[1846] 4. Emotion Recognition Integration
[1847] The server integrates a large-scale language model with an emotion recognition engine, which uses OpenCV and TensorFlow to analyze user emotions from audio and video data. The analysis results are reflected in the generated response.
[1848] 5. User Authentication and Response Generation
[1849] The terminal (smart glasses) performs personal authentication when the store clerk logs in. After logging in, when the store clerk interacts with the customer, voice input and video data are sent to the server. The server receives this data and passes it to an AI model and emotion recognition engine to generate an appropriate response.
[1850] 6. Real-time response
[1851] The server sends the generated response to the smart glasses in real time and displays it to the store clerk. This allows the store clerk to immediately provide an appropriate response to the customer. For example, if a customer asks, "I've been feeling stressed lately," the smart glasses will respond, "That's tough. What situations are making you feel stressed?"
[1852] 7. Reward Distribution
[1853] The server calculates and distributes rewards to data providers (store clerks and stores) based on the collected conversation data. Rewards are determined based on the frequency and quality of data usage.
[1854] Hardware and software used
[1855] Hardware: Smart glasses (e.g., Google Glass Enterprise Edition 2)
[1856] Server software: AWS Lambda, S3, EC2
[1857] Preprocessing and training: Python scripts, TensorFlow
[1858] Emotion recognition: OpenCV, TensorFlow
[1859] Authentication system: OAuth 2.0, Firebase Authentication
[1860] Examples of prompt statements
[1861] For example, if a customer asks, "I've been feeling stressed lately, especially at work...what should I do?" the system might respond with:
[1862] "That's tough. What situations are stressing you out?"
[1863] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1864] Step 1:
[1865] A user (a store clerk) wears smart glasses and serves customers in a store. The smart glasses collect audio and video data in real time and send it to a server. The input is conversation data (audio and video) between the customer and the store clerk, which is then sent to the server.
[1866] Step 2:
[1867] The server receives the transmitted conversation data and performs cleaning processes such as noise removal, filtering of inappropriate content and personal information. The data is then cleaned and stored in a database. The input is raw conversation data, and the output is cleaned data.
[1868] Step 3:
[1869] The server uses the cleaned conversation data to train a large-scale language model. It uses an AI framework such as TensorFlow to generate a model specialized for a specific conversation partner. The input is the cleaned conversation data, and the output is the trained large-scale language model.
[1870] Step 4:
[1871] The server integrates an emotion recognition engine with the trained large-scale language model. This engine uses OpenCV and TensorFlow to recognize user emotions from audio and video data. The input is audio and video data, and the output is recognized emotional information and response adjustments based on it.
[1872] Step 5:
[1873] The smart glasses at the terminal authenticate the store clerk by logging in to the system. During this process, the user's authentication information is sent to the server and verified. The input is the store clerk's authentication information, and the output is the authentication success or failure status.
[1874] Step 6:
[1875] The server receives real-time audio and video data when an authenticated store associate interacts with a customer and passes it to an AI model and emotion recognition engine. The AI model and emotion recognition engine then generate an appropriate response based on the user's input data. The input is real-time conversation data with the customer, and the output is the generated appropriate response.
[1876] Step 7:
[1877] The server sends the generated response to the smart glasses in real time and presents it to the store clerk. The store clerk responds to the customer based on the response. The input is the generated response, and the output is the response displayed on the smart glasses.
[1878] Step 8:
[1879] The server tracks all conversation sessions and calculates and distributes rewards to data providers based on the conversation data used. The input is the conversation session data, and the output is the reward calculation and distribution results.
[1880] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1881] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1882] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1883] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1884] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1885] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1886] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1887] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1888] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1889] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1890] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1891] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1892] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1893] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1894] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1895] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1896] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1897] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1898] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1899] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1900] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1901] The following is further disclosed regarding the above embodiment.
[1902] Understood. Below is a draft of the patent claims for the "Language Learning Talk Store" system, using high-level expressions.
[1903] (Claim 1)
[1904] a means for collecting conversation data;
[1905] A means of cleaning the collected conversation data and storing it in a database;
[1906] means for generating a large-scale language model that is trained using the stored conversational data;
[1907] a means for providing an appropriate AI model based on a particular conversation partner selected by the authenticated user;
[1908] A means for real-time conversation between the user and the AI model, with the AI model generating responses based on the user's input; and
[1909] means for calculating and distributing rewards to data providers based on the conversation data used;
[1910] A system including:
[1911] (Claim 2)
[1912] The system of claim 1, further comprising means for removing inappropriate content and personal information from the conversation data during pre-processing.
[1913] (Claim 3)
[1914] 10. The system of claim 1, further comprising means for providing a user interface for conversation data providers to upload conversation data.
[1915] "Example 1"
[1916] (Claim 1)
[1917] a means for collecting conversation data;
[1918] A means of cleaning the collected conversation data and storing it in a database;
[1919] means for generating a model that is trained using the stored conversation data;
[1920] a means for a user to select a particular conversation partner after being authenticated;
[1921] A means of conducting a conversation between the user and the model, and for the model to generate responses based on user input;
[1922] means for calculating and distributing rewards to data providers based on the conversation data used;
[1923] A system including:
[1924] (Claim 2)
[1925] The system of claim 1, further comprising means for removing inappropriate content and personal information from the conversation data during pre-processing.
[1926] (Claim 3)
[1927] 10. The system of claim 1, further comprising: means for providing an interface for conversation data providers to upload conversation data.
[1928] "Application Example 1"
[1929] (Claim 1)
[1930] a means for collecting conversation data;
[1931] A means of cleaning the collected conversation data and storing it in a database;
[1932] means for generating a large-scale language model that is trained using the stored conversational data;
[1933] a means for providing an appropriate AI model based on a particular conversation partner selected by the authenticated user;
[1934] A means for real-time conversation between the user and the AI model, with the AI model generating responses based on the user's input; and
[1935] means for calculating and distributing rewards to data providers based on the conversation data used;
[1936] means for providing an interface for a user to initiate an interactive simulation;
[1937] A means for tracking user conversation data and transferring rewards to conversation data providers through the system;
[1938] A system including:
[1939] (Claim 2)
[1940] The system of claim 1, further comprising means for removing inappropriate content and personal information from the conversation data during pre-processing.
[1941] (Claim 3)
[1942] 10. The system of claim 1, further comprising means for providing a user interface for conversation data providers to upload conversation data.
[1943] "Example 2: Combining Emotion Engines"
[1944] (Claim 1)
[1945] a means for collecting conversation data;
[1946] A means of cleaning the collected conversation data and storing it in a database;
[1947] means for generating a large-scale language model that is trained using the stored conversational data;
[1948] a means for providing an appropriate AI model based on a particular conversation partner selected by the authenticated user;
[1949] A means for real-time conversation between the user and the AI model, with the AI model generating responses based on the user's input; and
[1950] a means of integrating an emotion engine that recognizes user emotions and tailors responses;
[1951] means for calculating and distributing rewards to data providers based on the conversation data used;
[1952] A system including:
[1953] (Claim 2)
[1954] The system of claim 1, further comprising means for removing inappropriate content and personal information from the conversation data during pre-processing.
[1955] (Claim 3)
[1956] 10. The system of claim 1, further comprising means for providing a user interface for conversation data providers to upload conversation data.
[1957] "Application example 2 when combining emotion engines"
[1958] (Claim 1)
[1959] a means for collecting conversation data;
[1960] A means of cleaning the collected conversation data and storing it in a database;
[1961] means for generating a large-scale language model that is trained using the stored conversational data;
[1962] a means for providing an appropriate AI model based on a particular conversation partner selected by the authenticated user;
[1963] A means for real-time conversation between the user and the AI model, with the AI model generating responses based on the user's input; and
[1964] means for generating a response based on the user's emotion using an emotion recognition engine;
[1965] A means for analyzing the audio and video data collected using the smart glasses and providing an appropriate response to a store clerk in real time;
[1966] means for calculating and distributing rewards to data providers based on the conversation data used;
[1967] A system including:
[1968] (Claim 2)
[1969] The system of claim 1, further comprising means for removing inappropriate content and personal information from the conversation data during pre-processing.
[1970] (Claim 3)
[1971] 10. The system of claim 1, further comprising means for providing a user interface for conversation data providers to upload conversation data. [Explanation of symbols]
[1972] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. a means for collecting conversation data; A means of cleaning the collected conversation data and storing it in a database; means for generating a large-scale language model that is trained using the stored conversational data; a means for providing an appropriate AI model based on a particular conversation partner selected by the authenticated user; A means for real-time conversation between the user and the AI model, with the AI model generating responses based on the user's input; and means for calculating and distributing rewards to data providers based on the conversation data used; A system including:
2. The system of claim 1 further comprising means for removing inappropriate content and personal information from the conversation data during pre-processing.
3. 10. The system of claim 1, further comprising means for providing a user interface for conversation data providers to upload conversation data.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A