System
The system addresses language learning monotony by integrating beer brewing processes with generative AI to create interactive dialogue scenarios, enhancing learning engagement and cultural understanding.
Patent Information
- Application Number
- JP2024119136
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-24
- Publication Date
- 2026-02-05
Smart Images

Figure 2026018075000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In today's world, language learning requires sustained effort, and maintaining learning efficiency and motivation presents challenges. In particular, there are few opportunities to use the language learned in real life or to gain a deep understanding of its cultural background. Furthermore, the language learning process often feels monotonous, making it difficult to continue. Therefore, the present invention aims to simultaneously achieve language learning and intercultural understanding by using the beer brewing process, a topic of interest to users, thereby enhancing the enjoyment and effectiveness of learning. [Means for solving the problem]
[0005] The present invention provides a system including: means for receiving language selection information and basic information input by a user; means for generating a dialogue scenario related to the beer brewing process; means for providing the generated dialogue scenario in the user's language of choice; means for recognizing questions and responses input by the user and generating appropriate responses; and means for outputting the generated responses to the user. This system allows users to naturally deepen their language skills and intercultural understanding through dialogue-based learning on the subject of the beer brewing process. Furthermore, by providing a means for recording the user's learning progress and transmitting the recording to a server, learning support optimized for each individual user can be provided. Furthermore, by providing a means for periodically adding the latest information on the beer brewing process to a database and providing it to the user, a learning experience based on constantly new information can be provided.
[0006] The "user information selection means" is a function for inputting the language and basic information used by the user and receiving this information.
[0007] The "dialogue scenario generation means" is a function for creating a dialogue scenario relating to the beer brewing process and providing it to the user in a format suitable for the user.
[0008] The "language providing means" is a function for displaying or audibly outputting the generated dialogue scenario in a language selected by the user.
[0009] The "user question recognition means" is a function that recognizes questions and responses entered by the user and converts them into text.
[0010] The "appropriate response generating means" is a function for generating appropriate responses to user questions and responses and providing them as a system.
[0011] The "answer output means" is a function that outputs the generated answer to the user in voice or text format.
[0012] The "study progress recording means" is a function that records the user's learning progress and transmits that information to the server.
[0013] The "means for providing the latest information" is a function that periodically adds the latest dialogue scenarios and information about the beer brewing process to the database and provides them to the user. [Brief explanation of the drawings]
[0014] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0015] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0016] First, the terms used in the following description will be explained.
[0017] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0018] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0019] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0020] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0022] [First embodiment]
[0023] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0024] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0025] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0026] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0027] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0029] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0030] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0031] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0032] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0033] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0034] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0035] This invention relates to a system that allows users to deepen their language learning and intercultural understanding through the beer brewing process. The system involves three components: a server, a terminal, and a user, and clarifies the roles of each component.
[0036] Server Processing
[0037] The server serves as the core of this system and provides the following specific functions:
[0038] 1. Receiving user information:
[0039] The server receives language selection information and basic information input by the user through the terminal.
[0040] 2. Generating dialogue scenarios:
[0041] The server generates a dialogue scenario based on the beer brewing process. This dialogue scenario is structured appropriately based on the language selected by the user. Advanced generative AI such as ChatGPT is used to achieve natural dialogue.
[0042] 3. Response Generation:
[0043] The server receives questions and responses from users and uses AI technology to generate appropriate responses.
[0044] 4. Database Update:
[0045] The server periodically adds the latest information about the beer production process and new dialogue scenarios to the database and provides them to users.
[0046] Terminal handling
[0047] The terminal provides the following specific functions as an interface with the user:
[0048] 1. User Interface:
[0049] The terminal acts as a virtual tour guide, displaying each stage of the beer production process in the user's language of choice.
[0050] 2. Speech recognition and output:
[0051] The terminal converts the user's voice input into text and also synthesizes the text received from the server into speech for the user to hear.
[0052] 3. Data synchronization:
[0053] The device periodically synchronizes learning progress with the server, providing the user with the most up-to-date information.
[0054] User Behavior
[0055] The user performs the following specific actions through this system:
[0056] 1. Initial Settings and Selections:
[0057] Launch the app, select the language you want to learn, and enter your basic information.
[0058] 2. Selection of beer production process:
[0059] Within the app, users select the stage of the beer-making process they want to learn about (e.g., fermentation).
[0060] 3. Execute the dialogue:
[0061] The user interacts with the system at each stage of the process to learn. For example, in response to a question about the fermentation process, "Tell me how alcohol is produced during fermentation," the server responds, "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide."
[0062] Specific examples
[0063] Example: A user is learning about the beer fermentation process.
[0064] 1. Initial Setup:
[0065] The user launches the app, selects Japanese, and enters basic information. The device sends the information to the server, which stores it in a database and generates a learning plan optimized for the user.
[0066] 2. Process Selection:
[0067] The user selects "fermentation process" and the terminal sends this information to the server, which then generates a corresponding dialogue scenario and sends it to the terminal.
[0068] 3. Start a conversation:
[0069] The device starts the dialogue scenario by asking, "Do you know how alcohol is produced during the fermentation process?" The user responds, "I don't know. Please tell me." The device sends this to the server, which responds, "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide," which the device then relays in voice.
[0070] 4. Continuing the dialogue:
[0071] The user asks a further question (e.g., "What types of yeast are there?"). The device recognizes the question and sends it to the server, which generates an appropriate response, which the device then speaks to the user.
[0072] This system allows users to deepen their language learning and intercultural understanding through natural dialogue while using the beer brewing process as a subject, thereby increasing their motivation to learn.
[0073] The processing flow will be explained below.
[0074] Step 1:
[0075] The user launches the app. The device displays a startup screen, prompting the user to select a language and enter basic information.
[0076] Step 2:
[0077] The user inputs the desired language and basic information, which the terminal then sends to the server.
[0078] Step 3:
[0079] The server stores the received user information in a database, generates an optimal learning plan for the user, and sends it to the device.
[0080] Step 4:
[0081] Based on the learning plan received from the server, the device displays a selection screen for the beer brewing process, prompting the user to select the process they wish to learn.
[0082] Step 5:
[0083] The user selects "Fermentation process." The terminal transmits the user's selection to the server.
[0084] Step 6:
[0085] The server generates a dialogue scenario corresponding to the "fermentation process." It uses generative AI such as ChatGPT to build the dialogue scenario in the language selected by the user.
[0086] Step 7:
[0087] The server transmits the generated dialogue scenario to the terminal.
[0088] Step 8:
[0089] The device displays a dialogue scenario to the user and also starts voice output, asking questions such as, "Do you know how alcohol is produced during the fermentation process?"
[0090] Step 9:
[0091] The user responds to the question (e.g., "I don't understand. Please tell me."). The device recognizes the user's voice, converts it into text, and sends it to the server.
[0092] Step 10:
[0093] The server generates an appropriate response based on the user's response, using AI techniques to create answers in natural language (e.g., "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide.").
[0094] Step 11:
[0095] The server sends the generated response to the terminal.
[0096] Step 12:
[0097] The terminal synthesizes the response from the server into voice and plays it back to the user.
[0098] Step 13:
[0099] The user continues to ask questions (e.g., "What kinds of yeast are there?"). The device converts the question into text using voice recognition and sends it to the server.
[0100] Step 14:
[0101] The server generates a response to the new question (e.g., "There are many different types of yeast, including ale yeast and lager yeast.") and sends it to the terminal.
[0102] Step 15:
[0103] The device synthesizes the response into voice and plays it back to the user. This process is repeated to advance learning in an interactive format.
[0104] Step 16:
[0105] The device periodically sends the user's learning progress to the server, which records the progress and uses it to optimize the learning content for the next session.
[0106] Step 17:
[0107] The server periodically updates the database, adding new dialogue scenarios and the latest information on beer production, and the terminal provides the new information to the user.
[0108] Example 1
[0109] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0110] Conventional language learning systems struggle to sustain users' interest and attention, leading to a decline in motivation to learn. Furthermore, most language learning systems simply focus on language acquisition, failing to incorporate cultural background and practical knowledge, making it difficult to deepen intercultural understanding. Furthermore, responses to user questions are mechanical, making it difficult to achieve natural dialogue. This tends to reduce learning efficiency and user satisfaction.
[0111] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0112] In this invention, the server includes means for receiving language selection information and basic information input by the user, means for generating a dialogue scenario related to a manufacturing process such as a fermentation process, means for providing the generated dialogue scenario in the user's selected language, means for recognizing questions and responses input by the user and generating appropriate replies, and means for outputting the generated replies to the user. This makes it possible to deepen language learning and intercultural understanding in a natural dialogue format while maintaining the user's interest, thereby improving learning efficiency and user satisfaction.
[0113] "User" refers to an individual who uses this system to input information and interact in order to deepen language learning and intercultural understanding.
[0114] "Language selection information" is information that a user inputs to the system when selecting the language they wish to learn.
[0115] "Basic information" refers to personal information such as name, age, and interests that a user provides to the system.
[0116] The "server" is the core of this system, and refers to a computer system that processes information received from the user and generates dialogue scenarios and responses.
[0117] The "manufacturing process" specifically refers to each stage of production activities, such as the fermentation process of beer, and dialogue scenarios are generated based on this theme.
[0118] A "dialogue scenario" is a planned conversation content for progressing learning and dialogue, which is generated based on the manufacturing process selected by the user.
[0119] "Response" refers to the words or documents generated by the system in response to a user's question or response.
[0120] "Speech recognition" is a technology that converts a user's voice input into text data.
[0121] "Speech synthesis" is a technology that converts text data into speech and reads it out to the user.
[0122] "Data Server" refers to the computer system used to record and store users' learning progress and latest information.
[0123] "Database" refers to a data storage system for systematically storing collected information and generated scenarios.
[0124] This invention is a system that allows users to deepen their language learning and intercultural understanding through the manufacturing process. This system involves three components: a server, a terminal, and a user, and each component clearly divides its role.
[0125] Server Processing
[0126] The server serves as the core of this system and provides the following specific functions:
[0127] 1. Receiving user information:
[0128] The server receives the language selection information and basic information entered by the user through the terminal, and records the information in a user profile database using specific software such as an HTTP server (e.g., Nginx) and a database server (e.g., MySQL).
[0129] 2. Generating dialogue scenarios:
[0130] The server generates a dialogue scenario based on a manufacturing process (e.g., a fermentation process). Here, generative AI (e.g., ChatGPT) is used to create a natural and interactive dialogue scenario. For example, the following prompts are used:
[0131] The user wants to learn about the "fermentation process." Generate a clear explanation of the fermentation process.
[0132] 3. Response Generation:
[0133] The server receives questions and responses from the user and uses generative AI to generate appropriate responses. For example, if the user asks, "Tell me how alcohol is produced during fermentation," the server sends the following prompt to the generative AI:
[0134] User: How is alcohol produced during fermentation?
[0135] The generative AI replies, "During the fermentation process, yeast converts sugars into alcohol and carbon dioxide."
[0136] 4. Database Update:
[0137] The server periodically adds the latest information about the manufacturing process and new interaction scenarios to the database and provides them to users. The latest information is collected through web crawlers and external APIs and added to the database management system (DBMS).
[0138] Terminal handling
[0139] The terminal provides the following specific functions as an interface with the user:
[0140] 1. User Interface:
[0141] The device acts as a virtual tour guide, displaying each stage of the manufacturing process in the user's language of choice, and uses cross-platform frameworks such as React Native for its software.
[0142] 2. Speech recognition and output:
[0143] The device uses speech recognition technology (e.g., Google Cloud Speech-to-Text API) to convert the user's voice input into text, and also converts the text received from the server into speech using speech synthesis technology (e.g., Amazon Polly) and provides it to the user.
[0144] 3. Data synchronization:
[0145] The device periodically synchronizes its learning progress with the server to provide the user with the latest information, using a long-lived connection over WebSocket or HTTP.
[0146] User Behavior
[0147] The user performs the following specific actions through this system:
[0148] 1. Initial Settings and Selections:
[0149] Users launch the app, select the language they want to learn, and enter basic information. The device sends this information to the server, which stores it in a database. The server then generates a personalized learning plan for the user.
[0150] 2. Selection of beer production process:
[0151] Within the app, the user selects the stage of the manufacturing process they want to learn about (e.g., the fermentation stage). The device sends this information to the server, which then generates a corresponding dialogue scenario and sends it to the device.
[0152] 3. Execute the dialogue:
[0153] The user interacts with the system at each stage of the process to learn: for example, if a user asks, "Tell me how alcohol is produced during fermentation," the server responds, "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide."
[0154] 4. Continuing the dialogue:
[0155] If the user asks a further question (e.g., "What types of yeast are there?"), the device recognizes this question and sends it to the server, which generates an appropriate response, which the device then speaks to the user.
[0156] This system allows users to deepen their language learning and intercultural understanding through a natural dialogue format while using manufacturing processes as a subject, thereby increasing their motivation to learn.
[0157] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0158] Step 1: User launches the app
[0159] Input: A user taps on a device to launch an app.
[0160] Output: A language selection screen will appear.
[0161] What happens: When a user launches the app, the device loads the app's initial screen. The app is built with a cross-platform framework such as React Native and prompts the user to select a language.
[0162] Step 2: User enters language and basic information
[0163] Input: Enter your language preference and basic information (name, age, interests, etc.).
[0164] Output: The entered information is sent to the server and stored in a database.
[0165] Specific operation: The user enters the language and basic information in the input form and presses the "Submit" button. The terminal sends this as an HTTP POST request to the server, which receives the request via Nginx and stores it in the MySQL database.
[0166] Step 3: The server generates the lesson plan
[0167] Input: User's language preference and basic information.
[0168] Output: User-optimized learning plan data.
[0169] Specific operation: The server generates an optimal learning plan based on the received user information, referencing the data in the database. The generated learning plan is returned to the device in JSON format.
[0170] Step 4: User selects stage in manufacturing process
[0171] Input: The stage of the manufacturing process you want to learn about (e.g., fermentation process).
[0172] Output: Stage selection information is sent to the server.
[0173] Specific operation: When a user selects "Fermentation process" in the app, the device sends this information in JSON format to the server as an HTTP POST request. The server receives this information.
[0174] Step 5: The server generates the dialogue scenario
[0175] Input: Manufacturing process stage information.
[0176] Output: Dialogue scenario data.
[0177] Specific operation: The server sends prompts to generative AI models such as ChatGPT based on the manufacturing process. For example, it sends the following prompts:
[0178] The user wants to learn about the "fermentation process." Generate a clear explanation of the fermentation process.
[0179] The server receives the scenario generated by ChatGPT and returns it to the terminal in JSON format.
[0180] Step 6: The device presents the dialogue scenario
[0181] Input: Dialogue scenario data.
[0182] Output: Presents the dialogue scenario to the user in voice and text.
[0183] Specific operation: The device analyzes the received dialogue scenario and presents it to the user in both voice and text format. It converts the text into speech using speech synthesis technology (e.g., Amazon Polly).
[0184] Step 7: User enters question and response
[0185] Input: User voice or text input.
[0186] Output: The entered questions and responses are sent to the server in text format.
[0187] Specific operation: The user responds with something like "I don't understand. Please tell me." The device uses the Google Cloud Speech-to-Text API to convert the speech into text and send it to the server.
[0188] Step 8: The server generates a response
[0189] Input: The user's question and response.
[0190] Output: The generated response data.
[0191] How it works: The server receives the user's question and sends the following prompt to the generative AI (e.g. ChatGPT):
[0192] User: How is alcohol produced during fermentation?
[0193] The server receives the response generated by ChatGPT: "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide." and sends it to the terminal.
[0194] Step 9: The device presents a response
[0195] Input: The generated response data.
[0196] Output: Presents the response to the user in audio and text.
[0197] Specific operation: The device analyzes the response received from the server and uses speech synthesis technology to present it to the user in the form of voice and text.
[0198] This series of steps allows users to deepen their language learning and intercultural understanding through a natural, interactive way during the manufacturing process.
[0199] (Application example 1)
[0200] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0201] Conventional language learning systems have limited content that attracts users' interest, making it difficult to maintain motivation. In particular, they are limited to learning grammar and vocabulary, with few opportunities to learn cultural background or practical applications. Another problem is that the lack of interactive learning makes it difficult to improve practical language skills.
[0202] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0203] In this invention, the server includes means for receiving language selection information and basic information input by a user, means for generating a dialogue scenario related to the beer brewing process, means for providing the generated dialogue scenario in a language selected by the user, means for recognizing questions and responses input by the user and generating appropriate responses, means for outputting the generated responses to the user, means for visually representing the beer brewing process in a virtual environment, and means for the user to learn a language through dialogue in the virtual environment. This enables the user to deepen practical language learning and intercultural understanding through interactive dialogue in the virtual environment using the beer brewing process as a subject.
[0204] "Language selection information and basic information entered by the user" refers to the language used for learning that the user selects for the system and basic personal information.
[0205] The "means for generating a dialogue scenario" refers to a system component that has the function of automatically creating a dialogue scenario with the user based on information about the beer brewing process.
[0206] "Means for providing in a language selected by the user" refers to a system component that has the function of translating the generated dialogue scenario into a specific language selected by the user and presenting information in that language.
[0207] "Means for recognizing questions and responses and generating appropriate responses" refers to a system component that has the function of recognizing and analyzing questions and responses entered by users in voice or text format and automatically generating appropriate responses to them.
[0208] "Means for outputting the generated response to the user" refers to a system component that has the function of presenting the response generated by the system to the user in the form of voice or text.
[0209] "Means for visually representing the beer production process in a virtual environment" refers to a component of a system that uses virtual reality technology to visually recreate the beer production process and provide users with an interactive experience.
[0210] "Means for users to learn a language through interaction in a virtual environment" refers to a component of a system that has the function of allowing users to learn a language through interaction in an environment created using virtual reality technology.
[0211] This invention relates to a system that aims to deepen language learning and intercultural understanding through the beer brewing process. Specifically, it utilizes virtual reality technology and generative AI models to provide users with an interactive learning experience.
[0212] Server Processing
[0213] The server serves as the core of this system and provides the following specific functions:
[0214] 1. Receiving user information:
[0215] The server receives the language selection information and basic information input by the user through the terminal and stores them in a database.
[0216] 2. Dialogue scenario generation:
[0217] The server generates a dialogue scenario based on the beer brewing process, and the dialogue scenario is structured appropriately based on the language selected by the user. Advanced generative AI models such as ChatGPT are used to achieve natural dialogue.
[0218] 3. Response Generation:
[0219] The server receives questions and responses from users and uses generative AI technology to generate appropriate responses.
[0220] 4. Data synchronization:
[0221] The server records the user's learning progress and periodically adds the latest information and new dialogue scenarios to the database and provides them to the user.
[0222] Terminal handling
[0223] The terminal provides the following specific functions as an interface with the user:
[0224] 1. User Interface:
[0225] It acts as a virtual tour guide, showing each stage of the beer production process in the user's language of choice.
[0226] 2. Speech Recognition and Synthesis:
[0227] The system converts the user's voice input into text, synthesizes the text received from the server into speech, and plays it back to the user.
[0228] 3. Virtual Environment:
[0229] Users use smart glasses or a head-mounted display (HMD) to visually experience the beer production process in a virtual environment.
[0230] User Behavior
[0231] The user performs the following specific actions through this system:
[0232] 1. Initial Settings and Selections:
[0233] Launch the app, select the language you want to learn, and enter some basic information. The device then sends the information to the server, which stores it in a database.
[0234] 2. Selection of beer production process:
[0235] The user selects the stage of the beer production process (e.g., fermentation) they want to learn about in the virtual environment. The device sends this information to the server, which then generates a corresponding dialogue scenario.
[0236] 3. Execute the dialogue:
[0237] For example, in response to a question about the fermentation process, "Tell me how alcohol is produced during fermentation," the server responds, "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide."
[0238] Specific examples
[0239] Consider a case where a user is learning about the beer fermentation process.
[0240] 1. Initial Setup:
[0241] The user launches the app, selects Japanese, and enters basic information. The device sends the information to the server, which stores it in a database and generates a learning plan optimized for the user.
[0242] 2. Process Selection:
[0243] The user selects "fermentation process" and the terminal sends this information to the server, which then generates a corresponding dialogue scenario and sends it to the terminal.
[0244] 3. Start a conversation:
[0245] The device starts the dialogue scenario by asking, "Do you know how alcohol is produced during fermentation?" The user responds, "I don't know. Please tell me." The device sends this to the server, which replies, "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide," which the device then relays.
[0246] 4. Further dialogue:
[0247] The user then asks a further question (e.g., "What kinds of yeast are there?"). The device recognizes the question and sends it to the server, which generates an appropriate response, which the device then relays to the user via voice.
[0248] Prompt Sentence Examples
[0249] Below is an example of a prompt sentence to input to the generative AI model.
[0250] User Question: How is alcohol produced during fermentation?
[0251] Correct Response: During the fermentation process, yeast converts sugars into alcohol and carbon dioxide.
[0252] This system allows users to deepen their practical language learning and intercultural understanding through interactive dialogue in a virtual environment using the beer brewing process as a subject.
[0253] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0254] Step 1:
[0255] Entering and receiving user information
[0256] The user launches the application and inputs the language they want to learn and basic information through smart glasses or a head-mounted display (HMD). The device receives this information and sends it to the server. The input data is the user's language of choice and personal information, and the output is the user's basic information and learning plan, which are stored in a database. The server receives this data and stores it in a database.
[0257] Step 2:
[0258] Dialogue scenario generation
[0259] The server generates a dialogue scenario based on the user's selected language and information about the beer brewing process. It uses a generative AI model such as ChatGPT to create a scenario for natural dialogue. The input data are details of the beer brewing process and the user's selected language, and the output is the generated dialogue scenario. The server generates the dialogue scenario and sends it to the device.
[0260] Step 3:
[0261] Building a virtual environment
[0262] Based on the received dialogue scenario, the device creates a virtual environment in the smart glasses or head-mounted display (HMD) worn by the user. The input data is the dialogue scenario and virtual environment setting information sent from the server, and the output is a visual virtual environment provided to the user. The device displays each stage of the beer brewing process in a virtual space.
[0263] Step 4:
[0264] Starting a conversation
[0265] The user selects a specific stage of the beer brewing process in the virtual environment and begins a dialogue. The device sends this selection information to the server, which then prepares a corresponding dialogue scenario. The input data is the information about the stage selected by the user, and the output is the dialogue scenario prepared by the server. The device plays the scenario aloud and presents questions to the user.
[0266] Step 5:
[0267] Speech Recognition and Processing
[0268] The user responds to the question by voice. The device converts the user's voice into text data using voice recognition software and sends the text data to the server. The input data is the user's voice data, and the output is text data. The server receives this data and uses it in the next step.
[0269] Step 6:
[0270] Response Generation
[0271] The server uses a generative AI model to generate an appropriate response based on the received text data. The input data is the text data received from the user, and the output is the generated response text. The server generates this response and sends it to the device.
[0272] Step 7:
[0273] Response output
[0274] The terminal converts the response text received from the server into speech using speech synthesis software and presents the response to the user. The input data is the text data received from the server, and the output is a voice message provided to the user. As a specific example, if the user asks, "Please tell me how alcohol is produced during the fermentation process," the terminal will respond by voice, "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide."
[0275] Step 8:
[0276] Record your learning progress
[0277] The device records the user's learning progress and sends it to the server, which stores this information in a database and updates the user's optimized learning plan. The input data is the user's learning progress information, and the output is the updated learning information stored in the database.
[0278] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0279] This invention is a system that aims to help users deepen their language learning and intercultural understanding through the beer brewing process, and by incorporating an emotion engine that recognizes the user's emotions, it provides a more personalized learning experience. The system involves three parties: a server, a terminal, and a user, and clarifies the roles of each part.
[0280] Server Processing
[0281] The server serves as the core of this system and provides the following specific functions:
[0282] 1. Receiving user information:
[0283] The server receives language selection information and basic information input by the user through the terminal.
[0284] 2. Generating dialogue scenarios:
[0285] The server generates a dialogue scenario based on the beer brewing process. This dialogue scenario is structured appropriately based on the language selected by the user. Advanced generative AI such as ChatGPT is used to achieve natural dialogue.
[0286] 3. Response Generation:
[0287] The server receives questions and responses from users and uses AI technology to generate appropriate responses.
[0288] 4. Emotional Information Processing:
[0289] The server receives the user's emotional information sent from the terminal and adjusts the content and tone of the response to the user based on this information.
[0290] 5. Database Update:
[0291] The server periodically adds the latest information about the beer production process and new dialogue scenarios to the database and provides them to users.
[0292] Terminal handling
[0293] The terminal provides the following specific functions as an interface with the user:
[0294] 1. User Interface:
[0295] The terminal acts as a virtual tour guide, displaying each stage of the beer production process in the user's language of choice.
[0296] 2. Speech recognition and output:
[0297] The terminal converts the user's voice input into text and also synthesizes the text received from the server into speech for the user to hear.
[0298] 3. Emotion recognition:
[0299] The device uses voice input and facial recognition technology to recognize the user's emotions and sends that information to a server.
[0300] 4. Data synchronization:
[0301] The device periodically synchronizes learning progress with the server, providing the user with the most up-to-date information.
[0302] User Behavior
[0303] The user performs the following specific actions through this system:
[0304] 1. Initial Settings and Selections:
[0305] Launch the app, select the language you want to learn, and enter your basic information.
[0306] 2. Selection of beer production process:
[0307] Within the app, users select the stage of the beer-making process they want to learn about (e.g., fermentation).
[0308] 3. Execute the dialogue:
[0309] The user interacts with the system at each stage of the process to learn. For example, in response to a question about the fermentation process, "Tell me how alcohol is produced during fermentation," the server responds, "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide."
[0310] Specific examples
[0311] Example: A user is learning about the beer fermentation process.
[0312] 1. Initial Setup:
[0313] The user launches the app, selects Japanese, and enters basic information. The device sends the information to the server, which stores it in a database and generates a learning plan optimized for the user.
[0314] 2. Process Selection:
[0315] The user selects "fermentation process" and the terminal sends this information to the server, which then generates a corresponding dialogue scenario and sends it to the terminal.
[0316] 3. Start a conversation:
[0317] The device starts the dialogue scenario by asking, "Do you know how alcohol is produced during the fermentation process?" The user responds, "I don't know. Please tell me." The device sends this to the server, which responds, "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide," which the device then relays in voice.
[0318] 4. Emotion Recognition and Processing:
[0319] If the user shows an emotional reaction to a question (e.g., a questioning or surprised expression), the device recognizes that emotion and sends it to the server. The server then adjusts the tone and content of the response based on the emotional information, and generates an appropriate response and sends it to the device.
[0320] 5. Continuing the dialogue:
[0321] The user asks a further question (e.g., "What types of yeast are there?"). The device recognizes the question and sends it to the server, which generates an appropriate response, which the device then speaks to the user.
[0322] In this way, users can deepen their language learning and intercultural understanding through natural dialogue while learning about the beer brewing process. The introduction of an emotion engine provides responses that correspond to the user's emotions, creating a more personalized learning experience.
[0323] The processing flow will be explained below.
[0324] Step 1:
[0325] The user launches the app. The device displays a startup screen, prompting the user to select a language and enter basic information.
[0326] Step 2:
[0327] The user inputs the desired language and basic information, which the terminal then sends to the server.
[0328] Step 3:
[0329] The server stores the received user information in a database, generates an optimal learning plan for the user, and sends it to the device.
[0330] Step 4:
[0331] Based on the learning plan received from the server, the device displays a selection screen for the beer brewing process, prompting the user to select the process they wish to learn.
[0332] Step 5:
[0333] The user selects "Fermentation process." The terminal transmits the user's selection to the server.
[0334] Step 6:
[0335] The server generates a dialogue scenario corresponding to the "fermentation process." It uses generative AI such as ChatGPT to build the dialogue scenario in the language selected by the user.
[0336] Step 7:
[0337] The server transmits the generated dialogue scenario to the terminal.
[0338] Step 8:
[0339] The device displays the dialogue scenario to the user and also starts voice output, asking the question, "Do you know how alcohol is produced during the fermentation process?"
[0340] Step 9:
[0341] The user responds to the question (e.g., "I don't understand. Please tell me."). The device recognizes the user's voice, converts it into text, and sends it to the server.
[0342] Step 10:
[0343] The server generates an appropriate response based on the user's response, using AI techniques to create answers in natural language (e.g., "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide.").
[0344] Step 11:
[0345] The server sends the generated response to the terminal.
[0346] Step 12:
[0347] The terminal synthesizes the response from the server into voice and plays it back to the user.
[0348] Step 13:
[0349] The user asks a follow-up question (e.g., "What kinds of yeast are there?"). The device uses voice recognition to convert the question into text and sends it to the server.
[0350] Step 14:
[0351] The server generates a response to the new question (e.g., "There are many different types of yeast, including ale yeast and lager yeast.") and sends it to the terminal.
[0352] Step 15:
[0353] The device synthesizes the response into speech and plays it back to the user.
[0354] Step 16:
[0355] The emotion engine recognizes the tone of voice and facial expressions used during the user's questions and responses, and the device analyzes this and transmits the user's emotional information to the server.
[0356] Step 17:
[0357] Based on the received emotional information, the server adjusts the next response to match the tone and content of the user's feelings.
[0358] Step 18:
[0359] The device will provide users with responses tailored to their emotions, enabling more personalized interactions.
[0360] Step 19:
[0361] The device periodically sends the user's learning progress to the server, which records the progress and uses it to optimize the learning content for the next session.
[0362] Step 20:
[0363] The server periodically updates the database, adding new dialogue scenarios and the latest information on beer production, and the terminal provides the new information to the user.
[0364] Example 2
[0365] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0366] Traditional language learning systems struggled to provide a personalized experience based on users' interests and emotions. They also lacked dynamic learning content that incorporated the latest information. As a result, it was difficult to maintain learners' motivation, and effective learning could not be expected.
[0367] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0368] In this invention, the server includes means for receiving language selection information and basic information input by a user, means for generating a dialogue scenario related to the beer brewing process, means for providing the generated dialogue scenario in the user's language of choice, means for recognizing questions and responses input by the user and generating appropriate replies, means for analyzing the user's emotional state and adjusting the content and tone of the replies, means for outputting the generated replies to the user, means for adding the latest information to the database, and means for periodically updating the information in the database and providing it to the user. This provides a personalized learning experience based on the user's interests and emotions, and enables dynamic learning incorporating the latest information, thereby enabling effective language learning.
[0369] The "means for receiving language selection information and basic information input by the user" refers to an interface for receiving the language selected by the user through the application and basic information such as name, language level, and learning purpose.
[0370] The "means for generating dialogue scenarios related to the beer production process" is a system that uses a generative AI model or the like to generate dialogue-style scenarios based on each stage of beer production.
[0371] The "means for providing the generated dialogue scenario in a language selected by the user" refers to a means for displaying or audibly outputting the generated dialogue scenario in a language selected by the user.
[0372] "Means for recognizing questions and responses entered by users and generating appropriate responses" refers to a system that utilizes AI technology to recognize questions and responses from users in text or voice and generate responses to them.
[0373] "Means for analyzing the user's emotional state and adjusting the content and tone of the response" refers to means for detecting the user's emotional state using voice analysis or facial recognition technology and adjusting the content and tone of the response according to that emotion.
[0374] The "means for outputting the generated response to the user" refers to a means for providing the generated response from the server to the user in the form of voice or text.
[0375] The "means for adding the latest information to the database" is a means for periodically collecting the latest information on the beer production process and adding it to the database.
[0376] "Means for periodically updating the information in the database and providing it to users" refers to means for keeping the information in the database up to date and providing that information to users.
[0377] This invention is a system that aims to help users deepen their language learning and intercultural understanding through the beer brewing process, and by incorporating an emotion engine that recognizes the user's emotions, it provides a more personalized learning experience. The system involves three parties: a server, a terminal, and a user, and clarifies the roles of each part.
[0378] Specific server processing
[0379] The server serves as the core of this system and provides the following specific functions:
[0380] 1. Receiving user information:
[0381] The server receives language selection information and basic information entered by the user through the terminal, which is stored in a database and used to create a user profile.
[0382] 2. Generating dialogue scenarios:
[0383] The server generates a dialogue scenario based on the beer brewing process. This dialogue scenario is structured appropriately based on the language selected by the user. Natural dialogue is achieved using generative AI models such as ChatGPT.
[0384] 3. Response Generation:
[0385] The server receives questions and responses from users and uses AI technology to generate appropriate responses, allowing users to learn specialized knowledge about beer brewing.
[0386] 4. Emotional Information Processing:
[0387] The server receives the user's emotional information sent from the terminal and adjusts the content and tone of the response to the user based on this information.
[0388] 5. Database Update:
[0389] The server periodically adds the latest information about the beer production process and new dialogue scenarios to the database and provides them to users.
[0390] Specific processing of the terminal
[0391] The terminal provides the following specific functions as an interface with the user:
[0392] 1. User Interface:
[0393] The terminal acts as a virtual tour guide, displaying each stage of the beer production process in the user's language of choice.
[0394] 2. Speech recognition and output:
[0395] The terminal converts the user's voice input into text and also synthesizes the text received from the server into speech for the user to hear.
[0396] 3. Emotion recognition:
[0397] The device uses voice input and facial recognition technology to recognize the user's emotions and sends that information to a server.
[0398] 4. Data synchronization:
[0399] The device periodically synchronizes learning progress with the server, providing the user with the most up-to-date information.
[0400] Specific user actions
[0401] The user will perform the following specific actions through this system:
[0402] 1. Initial Settings and Selections:
[0403] Users launch the app, select the language they want to learn, and enter basic information. The device sends this information to the server, which stores it in a database and generates a personalized learning plan.
[0404] 2. Selection of beer production process:
[0405] Within the app, users select the stage of the beer brewing process they want to learn about (e.g., the fermentation stage). The device sends this information to the server, which then generates a corresponding dialogue scenario and sends it to the device.
[0406] 3. Execute the dialogue:
[0407] The user interacts with the system at each stage of the process. For example, in response to a question about the fermentation process, "Tell me how alcohol is produced during fermentation," the server responds with, "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide." For example, the user can ask a question using the following prompt:
[0408] "Please explain the fermentation process of beer in detail."
[0409] What types of yeast are there?
[0410] summary
[0411] The system is designed to enable users to deepen their language learning and intercultural understanding through natural dialogue while learning about the beer brewing process. The system incorporates an emotion engine to provide responses based on the user's emotions, creating a more personalized learning experience. The system also adds the latest information to the database and provides it to users, ensuring they are constantly learning new knowledge.
[0412] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0413] Program processing flow
[0414] Step 1: Enter and submit user information
[0415] Specific behavior:
[0416] 1. Input: The user launches the app, selects Japanese or another language they want to learn, and enters basic information such as their name, language level, and learning purpose.
[0417] 2. Data processing: The device collects this information and converts it into the appropriate format.
[0418] 3. Send: The device creates and sends an API request to send the formatted information to the server.
[0419] 4. Output: The server saves the user's basic information and language preference in a database and creates a user profile.
[0420] Step 2: Generate dialogue scenarios
[0421] Specific behavior:
[0422] 1. Input: The server retrieves information about the beer production process from a database.
[0423] 2. Data processing: Based on the information obtained, the server converts the prompt text into a format suitable for sending to the generative AI model.
[0424] 3. Generation: A generative AI model like ChatGPT generates a dialogue scenario based on the prompt.
[0425] 4. Output: The server receives the generated dialogue scenario, converts it into an appropriate format, and sends it to the terminal to be presented in the language selected by the user.
[0426] Step 3: Handling User Questions and Responses
[0427] Specific behavior:
[0428] 1. Input: The user enters questions and responses by voice or text at each stage of the process, for example, "Tell me how alcohol is produced during fermentation."
[0429] 2. Data processing: The device uses voice recognition technology to convert the voice into text and sends the text data to the server.
[0430] 3. Generation: The server analyzes the user's question, sends it as a prompt to the generative AI model, and generates an appropriate response.
[0431] 4. Output: The server receives the generated response and sends it to the terminal, which then provides the response to the user using voice synthesis technology.
[0432] Step 4: Recognizing and processing emotional information
[0433] Specific behavior:
[0434] 1. Input: The user asks questions and responds to the device, and emotions are input from facial expressions and tone of voice.
[0435] 2. Data processing: The device uses voice analysis and facial recognition technology to analyze the user's emotional state, for example, detecting expressions of doubt or surprise.
[0436] 3. Send: The device sends the emotion analysis results to the server.
[0437] 4. Processing: The server adjusts the tone and content of the response based on the emotional information, generates an appropriate response, and sends it to the device.
[0438] 5. Output: The device communicates the adjusted response received from the server to the user via voice synthesis.
[0439] Step 5: Update the database
[0440] Specific behavior:
[0441] 1. Input: The server retrieves up-to-date information about the beer production process from an external resource.
[0442] 2. Data processing: The server converts the acquired information into a format for adding to the database.
[0443] 3. Add: The server adds new information to the database and updates existing information.
[0444] 4. Output: Based on the updated information, the server generates a new dialogue scenario and sends it to the terminal to provide the user with up-to-date information.
[0445] (Application example 2)
[0446] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0447] The purpose of this invention is to provide a language learning system based on the beer brewing process that can personalize the user's learning experience and provide an interactive and effective learning environment. In particular, the invention aims to deepen the user's interest and understanding in learning by recognizing the user's emotions and adjusting the content and tone of responses accordingly, and to provide an experience based on actual procedures and culture by visually displaying the beer brewing process in a virtual environment.
[0448] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving language selection information and basic information input by the user, means for generating a dialogue scenario related to the beer brewing process, means for providing the generated dialogue scenario in the user's language, means for recognizing questions and responses input by the user and generating appropriate replies, means for outputting the generated replies to the user, means for recognizing emotions and adjusting the content and tone of the responses, and means for visually displaying the beer brewing process in a virtual environment. This enables users to deepen their language learning and intercultural understanding through a natural dialogue format through the beer brewing process. Furthermore, the use of emotion recognition and a virtual environment makes the learning experience more interactive and personalized.
[0449] "Means for receiving user-entered language selection information and basic information" refers to a device for entering the user's preferred language and typical profile information (e.g., name, age, affiliation, etc.) and the function for transmitting that information to the system.
[0450] "Means for generating dialogue scenarios related to the beer production process" refers to a function for generating scenarios for dialogue in a language selected by the learner based on each stage of beer production.
[0451] "Means for providing the generated dialogue scenario in a language selected by the user" refers to an interface and device for translating the generated dialogue scenario into a language selected by the user and effectively providing it.
[0452] "Means of recognizing questions and responses entered by the user and generating appropriate responses" refers to a function that uses voice recognition technology and natural language processing technology to understand the content entered by the user and automatically generate appropriate responses based on that.
[0453] "Means for outputting the generated response to the user" refers to an output device or software that conveys the system-generated response to the user in audio or text form.
[0454] "Means for recognizing emotions and adjusting the content and tone of responses" refers to a function that analyzes emotions from the user's facial expressions and voice, and optimizes the system's response content and speaking style according to those emotions.
[0455] "Means for visually displaying the beer production process in a virtual environment" refers to a function that uses virtual reality or augmented reality technology to display the production process so that the user can visually understand each stage of beer production.
[0456] This invention is a system that aims to help users deepen their language learning and intercultural understanding through the beer brewing process, and by incorporating an emotion engine that recognizes the user's emotions, it provides a more personalized learning experience. Here, we will explain in detail how to specifically implement the invention.
[0457] System Overview
[0458] This system is realized by the collaboration of three parties: the server, the terminal, and the user, each fulfilling their respective roles. The server is responsible for all main processing, while the terminal functions as an interface with the user, allowing the user to learn through dialogue.
[0459] Program processing
[0460] Server Processing
[0461] 1. Receiving user information:
[0462] The server receives the language selection information and basic information entered by the user through the terminal, and this function is used to securely store the entered data and generate the necessary study plan.
[0463] 2. Generating dialogue scenarios:
[0464] The server generates a dialogue scenario based on the beer brewing process. This dialogue scenario is structured appropriately based on the language selected by the user. It uses OpenAI's generative AI model to achieve natural dialogue.
[0465] 3. Response Generation:
[0466] The server receives questions and responses from users and uses natural language processing technology to generate appropriate responses, using the latest generative AI models.
[0467] 4. Emotional Information Processing:
[0468] The server receives the user's emotional information sent from the terminal and adjusts the content and tone of the response to the user based on this information.
[0469] 5. Virtual environment display:
[0470] The server generates data to visually display each stage of the beer production process and sends it to the device, allowing users to experience it in a realistic way through visual content.
[0471] Terminal handling
[0472] 1. User Interface:
[0473] The terminal acts as a virtual tour guide, displaying each stage of the beer production process in the user's language of choice, with an intuitive and easy-to-use interface.
[0474] 2. Speech recognition and output:
[0475] The device converts the user's voice input into text and synthesizes the text received from the server into speech for the user to hear. The libraries used are "speech_recognition" and "pyttsx3".
[0476] 3. Emotion recognition:
[0477] The device recognizes the user's emotions using voice input and facial recognition technology, and sends the information to the server. Emotion recognition is performed using OpenCV and a pre-trained emotion recognition model.
[0478] 4. Data synchronization:
[0479] The device periodically synchronizes learning progress with the server, providing the user with the most up-to-date information.
[0480] User Behavior
[0481] 1. Initial Settings and Selections:
[0482] Users launch the app, select the language they want to learn, and enter basic information, which the device then sends to the server.
[0483] 2. Selection of beer production process:
[0484] Within the app, users select the stage of the beer-making process they want to learn about (e.g., fermentation), and the device sends this information to the server, which then generates a corresponding dialogue scenario.
[0485] 3. Execute the dialogue:
[0486] The user interacts with the system at each stage of the process to learn. For example, in response to a question about the fermentation process, "Tell me how alcohol is produced during fermentation," the server responds, "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide."
[0487] 4. Emotion Recognition and Processing:
[0488] If the user shows an emotional reaction to a question (e.g., a questioning or surprised expression), the device recognizes that emotion and transmits it to the server. The server then adjusts the tone and content of the response based on the emotional information, and generates an appropriate response and transmits it to the device.
[0489] 5. Continuing the dialogue:
[0490] If the user asks a further question (e.g., "What types of yeast are there?"), the device recognizes the question and sends it to the server, which generates an appropriate response, which the device then speaks to the user.
[0491] Prompt Sentence Examples
[0492] "A Guide to the Beer Making Process: Tell Me About Fermentation."
[0493] "A Guide to the Beer Making Process: Explain the Malting Steps."
[0494] In this way, users can deepen their language learning and intercultural understanding through natural dialogue while learning about the beer brewing process. The introduction of an emotion engine provides responses that correspond to the user's emotions, creating a more personalized learning experience.
[0495] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0496] Step 1:
[0497] The server receives language selection information and basic information entered by the user through the device, including profile information such as the user's preferred language, name, and age. The received information is stored in a database and used to generate a learning plan.
[0498] Step 2:
[0499] The server generates a dialogue scenario about the beer brewing process. It uses a generative AI model to create a natural dialogue scenario based on the language selected by the user. This dialogue scenario contains content related to the stage of the process the user wants to learn about (e.g., fermentation). The input used here is the user information saved in step 1, and the output is the generated dialogue scenario.
[0500] Step 3:
[0501] The terminal provides the generated dialogue scenario in the language selected by the user. The user interface (UI) displays the dialogue scenario visually and audibly, allowing the user to begin the dialogue. The input here is the dialogue scenario sent from the server, and the output is the dialogue scenario displayed on the user interface.
[0502] Step 4:
[0503] The user inputs a question or response into the terminal. Speech recognition technology is used to convert the user's voice input into text, and the text is sent to the server. The input is the user's voice, and the output is the user's question or response converted into text.
[0504] Step 5:
[0505] The server recognizes the questions and responses entered by the user and generates an appropriate response. It uses a generative AI model to perform natural language processing and generate an appropriate response to the user's question. The input is the text received in step 4, and the output is the generated response.
[0506] Step 6:
[0507] The terminal outputs the response received from the server to the user as voice. It uses speech synthesis technology (pyttsx3) to convey information in a way that is easy for the user to understand. The input is the text of the response sent from the server, and the output is the voice response.
[0508] Step 7:
[0509] The device recognizes the user's emotions and sends the information to the server. Emotion recognition uses OpenCV and a pre-trained emotion recognition model. The input is the user's facial image and voice, and the output is the recognized emotion information.
[0510] Step 8:
[0511] The server adjusts the content and tone of the response based on the emotional information. It changes the tone and content of the response depending on the emotion expressed by the user, generating a more personalized response. The input is the emotional information received in step 7, and the output is the adjusted response.
[0512] Step 9:
[0513] The device outputs the adjusted response to the user as voice, again using speech synthesis technology to provide an adapted response to the user. The input is the adjusted response text sent from the server, and the output is the voice response.
[0514] Step 10:
[0515] The terminal visually displays the beer brewing process in a virtual environment. Using virtual reality and augmented reality technology, the user can visually experience the beer brewing process. The input is visual data sent from the server, and the output is the beer brewing process displayed in the virtual environment.
[0516] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0517] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0518] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0519] [Second embodiment]
[0520] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0521] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0522] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0523] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0524] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0525] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0526] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0527] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0528] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0529] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0530] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0531] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0532] This invention relates to a system that allows users to deepen their language learning and intercultural understanding through the beer brewing process. The system involves three components: a server, a terminal, and a user, and clarifies the roles of each component.
[0533] Server Processing
[0534] The server serves as the core of this system and provides the following specific functions:
[0535] 1. Receiving user information:
[0536] The server receives language selection information and basic information input by the user through the terminal.
[0537] 2. Generating dialogue scenarios:
[0538] The server generates a dialogue scenario based on the beer brewing process. This dialogue scenario is structured appropriately based on the language selected by the user. Advanced generative AI such as ChatGPT is used to achieve natural dialogue.
[0539] 3. Response Generation:
[0540] The server receives questions and responses from users and uses AI technology to generate appropriate responses.
[0541] 4. Database Update:
[0542] The server periodically adds the latest information about the beer production process and new dialogue scenarios to the database and provides them to users.
[0543] Terminal handling
[0544] The terminal provides the following specific functions as an interface with the user:
[0545] 1. User Interface:
[0546] The terminal acts as a virtual tour guide, displaying each stage of the beer production process in the user's language of choice.
[0547] 2. Speech recognition and output:
[0548] The terminal converts the user's voice input into text and also synthesizes the text received from the server into speech for the user to hear.
[0549] 3. Data synchronization:
[0550] The device periodically synchronizes learning progress with the server, providing the user with the most up-to-date information.
[0551] User Behavior
[0552] The user performs the following specific actions through this system:
[0553] 1. Initial Settings and Selections:
[0554] Launch the app, select the language you want to learn, and enter your basic information.
[0555] 2. Selection of beer production process:
[0556] Within the app, users select the stage of the beer-making process they want to learn about (e.g., fermentation).
[0557] 3. Execute the dialogue:
[0558] The user interacts with the system at each stage of the process to learn. For example, in response to a question about the fermentation process, "Tell me how alcohol is produced during fermentation," the server responds, "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide."
[0559] Specific examples
[0560] Example: A user is learning about the beer fermentation process.
[0561] 1. Initial Setup:
[0562] The user launches the app, selects Japanese, and enters basic information. The device sends the information to the server, which stores it in a database and generates a learning plan optimized for the user.
[0563] 2. Process Selection:
[0564] The user selects "fermentation process" and the terminal sends this information to the server, which then generates a corresponding dialogue scenario and sends it to the terminal.
[0565] 3. Start a conversation:
[0566] The device starts the dialogue scenario by asking, "Do you know how alcohol is produced during the fermentation process?" The user responds, "I don't know. Please tell me." The device sends this to the server, which responds, "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide," which the device then relays in voice.
[0567] 4. Continuing the dialogue:
[0568] The user asks a further question (e.g., "What types of yeast are there?"). The device recognizes the question and sends it to the server, which generates an appropriate response, which the device then speaks to the user.
[0569] This system allows users to deepen their language learning and intercultural understanding through natural dialogue while using the beer brewing process as a subject, thereby increasing their motivation to learn.
[0570] The processing flow will be explained below.
[0571] Step 1:
[0572] The user launches the app. The device displays a startup screen, prompting the user to select a language and enter basic information.
[0573] Step 2:
[0574] The user inputs the desired language and basic information, which the terminal then sends to the server.
[0575] Step 3:
[0576] The server stores the received user information in a database, generates an optimal learning plan for the user, and sends it to the device.
[0577] Step 4:
[0578] Based on the learning plan received from the server, the device displays a selection screen for the beer brewing process, prompting the user to select the process they wish to learn.
[0579] Step 5:
[0580] The user selects "Fermentation process." The terminal transmits the user's selection to the server.
[0581] Step 6:
[0582] The server generates a dialogue scenario corresponding to the "fermentation process." It uses generative AI such as ChatGPT to build the dialogue scenario in the language selected by the user.
[0583] Step 7:
[0584] The server transmits the generated dialogue scenario to the terminal.
[0585] Step 8:
[0586] The device displays a dialogue scenario to the user and also starts voice output, asking questions such as, "Do you know how alcohol is produced during the fermentation process?"
[0587] Step 9:
[0588] The user responds to the question (e.g., "I don't understand. Please tell me."). The device recognizes the user's voice, converts it into text, and sends it to the server.
[0589] Step 10:
[0590] The server generates an appropriate response based on the user's response, using AI techniques to create answers in natural language (e.g., "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide.").
[0591] Step 11:
[0592] The server sends the generated response to the terminal.
[0593] Step 12:
[0594] The terminal synthesizes the response from the server into voice and plays it back to the user.
[0595] Step 13:
[0596] The user continues to ask questions (e.g., "What kinds of yeast are there?"). The device converts the question into text using voice recognition and sends it to the server.
[0597] Step 14:
[0598] The server generates a response to the new question (e.g., "There are many different types of yeast, including ale yeast and lager yeast.") and sends it to the terminal.
[0599] Step 15:
[0600] The device synthesizes the response into voice and plays it back to the user. This process is repeated to advance learning in an interactive format.
[0601] Step 16:
[0602] The device periodically sends the user's learning progress to the server, which records the progress and uses it to optimize the learning content for the next session.
[0603] Step 17:
[0604] The server periodically updates the database, adding new dialogue scenarios and the latest information on beer production, and the terminal provides the new information to the user.
[0605] Example 1
[0606] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0607] Conventional language learning systems struggle to sustain users' interest and attention, leading to a decline in motivation to learn. Furthermore, most language learning systems simply focus on language acquisition, failing to incorporate cultural background and practical knowledge, making it difficult to deepen intercultural understanding. Furthermore, responses to user questions are mechanical, making it difficult to achieve natural dialogue. This tends to reduce learning efficiency and user satisfaction.
[0608] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0609] In this invention, the server includes means for receiving language selection information and basic information input by the user, means for generating a dialogue scenario related to a manufacturing process such as a fermentation process, means for providing the generated dialogue scenario in the user's selected language, means for recognizing questions and responses input by the user and generating appropriate replies, and means for outputting the generated replies to the user. This makes it possible to deepen language learning and intercultural understanding in a natural dialogue format while maintaining the user's interest, thereby improving learning efficiency and user satisfaction.
[0610] "User" refers to an individual who uses this system to input information and interact in order to deepen language learning and intercultural understanding.
[0611] "Language selection information" is information that a user inputs to the system when selecting the language they wish to learn.
[0612] "Basic information" refers to personal information such as name, age, and interests that a user provides to the system.
[0613] The "server" is the core of this system, and refers to a computer system that processes information received from the user and generates dialogue scenarios and responses.
[0614] The "manufacturing process" specifically refers to each stage of production activities, such as the fermentation process of beer, and dialogue scenarios are generated based on this theme.
[0615] A "dialogue scenario" is a planned conversation content for progressing learning and dialogue, which is generated based on the manufacturing process selected by the user.
[0616] "Response" refers to the words or documents generated by the system in response to a user's question or response.
[0617] "Speech recognition" is a technology that converts a user's voice input into text data.
[0618] "Speech synthesis" is a technology that converts text data into speech and reads it out to the user.
[0619] "Data Server" refers to the computer system used to record and store users' learning progress and latest information.
[0620] "Database" refers to a data storage system for systematically storing collected information and generated scenarios.
[0621] This invention is a system that allows users to deepen their language learning and intercultural understanding through the manufacturing process. This system involves three components: a server, a terminal, and a user, and each component clearly divides its role.
[0622] Server Processing
[0623] The server serves as the core of this system and provides the following specific functions:
[0624] 1. Receiving user information:
[0625] The server receives the language selection information and basic information entered by the user through the terminal, and records the information in a user profile database using specific software such as an HTTP server (e.g., Nginx) and a database server (e.g., MySQL).
[0626] 2. Generating dialogue scenarios:
[0627] The server generates a dialogue scenario based on a manufacturing process (e.g., a fermentation process). Here, generative AI (e.g., ChatGPT) is used to create a natural and interactive dialogue scenario. For example, the following prompts are used:
[0628] The user wants to learn about the "fermentation process." Generate a clear explanation of the fermentation process.
[0629] 3. Response Generation:
[0630] The server receives questions and responses from the user and uses generative AI to generate appropriate responses. For example, if the user asks, "Tell me how alcohol is produced during fermentation," the server sends the following prompt to the generative AI:
[0631] User: How is alcohol produced during fermentation?
[0632] The generative AI replies, "During the fermentation process, yeast converts sugars into alcohol and carbon dioxide."
[0633] 4. Database Update:
[0634] The server periodically adds the latest information about the manufacturing process and new interaction scenarios to the database and provides them to users. The latest information is collected through web crawlers and external APIs and added to the database management system (DBMS).
[0635] Terminal handling
[0636] The terminal provides the following specific functions as an interface with the user:
[0637] 1. User Interface:
[0638] The device acts as a virtual tour guide, displaying each stage of the manufacturing process in the user's language of choice, and uses cross-platform frameworks such as React Native for its software.
[0639] 2. Speech recognition and output:
[0640] The device uses speech recognition technology (e.g., Google Cloud Speech-to-Text API) to convert the user's voice input into text, and also converts the text received from the server into speech using speech synthesis technology (e.g., Amazon Polly) and provides it to the user.
[0641] 3. Data synchronization:
[0642] The device periodically synchronizes its learning progress with the server to provide the user with the latest information, using a long-lived connection over WebSocket or HTTP.
[0643] User Behavior
[0644] The user performs the following specific actions through this system:
[0645] 1. Initial Settings and Selections:
[0646] Users launch the app, select the language they want to learn, and enter basic information. The device sends this information to the server, which stores it in a database. The server then generates a personalized learning plan for the user.
[0647] 2. Selection of beer production process:
[0648] Within the app, the user selects the stage of the manufacturing process they want to learn about (e.g., the fermentation stage). The device sends this information to the server, which then generates a corresponding dialogue scenario and sends it to the device.
[0649] 3. Execute the dialogue:
[0650] The user interacts with the system at each stage of the process to learn: for example, if a user asks, "Tell me how alcohol is produced during fermentation," the server responds, "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide."
[0651] 4. Continuing the dialogue:
[0652] If the user asks a further question (e.g., "What types of yeast are there?"), the device recognizes this question and sends it to the server, which generates an appropriate response, which the device then speaks to the user.
[0653] This system allows users to deepen their language learning and intercultural understanding through a natural dialogue format while using manufacturing processes as a subject, thereby increasing their motivation to learn.
[0654] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0655] Step 1: User launches the app
[0656] Input: A user taps on a device to launch an app.
[0657] Output: A language selection screen will appear.
[0658] What happens: When a user launches the app, the device loads the app's initial screen. The app is built with a cross-platform framework such as React Native and prompts the user to select a language.
[0659] Step 2: User enters language and basic information
[0660] Input: Enter your language preference and basic information (name, age, interests, etc.).
[0661] Output: The entered information is sent to the server and stored in a database.
[0662] Specific operation: The user enters the language and basic information in the input form and presses the "Submit" button. The terminal sends this as an HTTP POST request to the server, which receives the request via Nginx and stores it in the MySQL database.
[0663] Step 3: The server generates the lesson plan
[0664] Input: User's language preference and basic information.
[0665] Output: User-optimized learning plan data.
[0666] Specific operation: The server generates an optimal learning plan based on the received user information, referencing the data in the database. The generated learning plan is returned to the device in JSON format.
[0667] Step 4: User selects stage in manufacturing process
[0668] Input: The stage of the manufacturing process you want to learn about (e.g., fermentation process).
[0669] Output: Stage selection information is sent to the server.
[0670] Specific operation: When a user selects "Fermentation process" in the app, the device sends this information in JSON format to the server as an HTTP POST request. The server receives this information.
[0671] Step 5: The server generates the dialogue scenario
[0672] Input: Manufacturing process stage information.
[0673] Output: Dialogue scenario data.
[0674] Specific operation: The server sends prompts to generative AI models such as ChatGPT based on the manufacturing process. For example, it sends the following prompts:
[0675] The user wants to learn about the "fermentation process." Generate a clear explanation of the fermentation process.
[0676] The server receives the scenario generated by ChatGPT and returns it to the terminal in JSON format.
[0677] Step 6: The device presents the dialogue scenario
[0678] Input: Dialogue scenario data.
[0679] Output: Presents the dialogue scenario to the user in voice and text.
[0680] Specific operation: The device analyzes the received dialogue scenario and presents it to the user in both voice and text format. It converts the text into speech using speech synthesis technology (e.g., Amazon Polly).
[0681] Step 7: User enters question and response
[0682] Input: User voice or text input.
[0683] Output: The entered questions and responses are sent to the server in text format.
[0684] Specific operation: The user responds with something like "I don't understand. Please tell me." The device uses the Google Cloud Speech-to-Text API to convert the speech into text and send it to the server.
[0685] Step 8: The server generates a response
[0686] Input: The user's question and response.
[0687] Output: The generated response data.
[0688] How it works: The server receives the user's question and sends the following prompt to the generative AI (e.g. ChatGPT):
[0689] User: How is alcohol produced during fermentation?
[0690] The server receives the response generated by ChatGPT: "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide." and sends it to the terminal.
[0691] Step 9: The device presents a response
[0692] Input: The generated response data.
[0693] Output: Presents the response to the user in audio and text.
[0694] Specific operation: The device analyzes the response received from the server and uses speech synthesis technology to present it to the user in the form of voice and text.
[0695] This series of steps allows users to deepen their language learning and intercultural understanding through a natural, interactive way during the manufacturing process.
[0696] (Application example 1)
[0697] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0698] Conventional language learning systems have limited content that attracts users' interest, making it difficult to maintain motivation. In particular, they are limited to learning grammar and vocabulary, with few opportunities to learn cultural background or practical applications. Another problem is that the lack of interactive learning makes it difficult to improve practical language skills.
[0699] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0700] In this invention, the server includes means for receiving language selection information and basic information input by a user, means for generating a dialogue scenario related to the beer brewing process, means for providing the generated dialogue scenario in a language selected by the user, means for recognizing questions and responses input by the user and generating appropriate responses, means for outputting the generated responses to the user, means for visually representing the beer brewing process in a virtual environment, and means for the user to learn a language through dialogue in the virtual environment. This enables the user to deepen practical language learning and intercultural understanding through interactive dialogue in the virtual environment using the beer brewing process as a subject.
[0701] "Language selection information and basic information entered by the user" refers to the language used for learning that the user selects for the system and basic personal information.
[0702] The "means for generating a dialogue scenario" refers to a system component that has the function of automatically creating a dialogue scenario with the user based on information about the beer brewing process.
[0703] "Means for providing in a language selected by the user" refers to a system component that has the function of translating the generated dialogue scenario into a specific language selected by the user and presenting information in that language.
[0704] "Means for recognizing questions and responses and generating appropriate responses" refers to a system component that has the function of recognizing and analyzing questions and responses entered by users in voice or text format and automatically generating appropriate responses to them.
[0705] "Means for outputting the generated response to the user" refers to a system component that has the function of presenting the response generated by the system to the user in the form of voice or text.
[0706] "Means for visually representing the beer production process in a virtual environment" refers to a component of a system that uses virtual reality technology to visually recreate the beer production process and provide users with an interactive experience.
[0707] "Means for users to learn a language through interaction in a virtual environment" refers to a component of a system that has the function of allowing users to learn a language through interaction in an environment created using virtual reality technology.
[0708] This invention relates to a system that aims to deepen language learning and intercultural understanding through the beer brewing process. Specifically, it utilizes virtual reality technology and generative AI models to provide users with an interactive learning experience.
[0709] Server Processing
[0710] The server serves as the core of this system and provides the following specific functions:
[0711] 1. Receiving user information:
[0712] The server receives the language selection information and basic information input by the user through the terminal and stores them in a database.
[0713] 2. Dialogue scenario generation:
[0714] The server generates a dialogue scenario based on the beer brewing process, and the dialogue scenario is structured appropriately based on the language selected by the user. Advanced generative AI models such as ChatGPT are used to achieve natural dialogue.
[0715] 3. Response Generation:
[0716] The server receives questions and responses from users and uses generative AI technology to generate appropriate responses.
[0717] 4. Data synchronization:
[0718] The server records the user's learning progress and periodically adds the latest information and new dialogue scenarios to the database and provides them to the user.
[0719] Terminal handling
[0720] The terminal provides the following specific functions as an interface with the user:
[0721] 1. User Interface:
[0722] It acts as a virtual tour guide, showing each stage of the beer production process in the user's language of choice.
[0723] 2. Speech Recognition and Synthesis:
[0724] The system converts the user's voice input into text, synthesizes the text received from the server into speech, and plays it back to the user.
[0725] 3. Virtual Environment:
[0726] Users use smart glasses or a head-mounted display (HMD) to visually experience the beer production process in a virtual environment.
[0727] User Behavior
[0728] The user performs the following specific actions through this system:
[0729] 1. Initial Settings and Selections:
[0730] Launch the app, select the language you want to learn, and enter some basic information. The device then sends the information to the server, which stores it in a database.
[0731] 2. Selection of beer production process:
[0732] The user selects the stage of the beer production process (e.g., fermentation) they want to learn about in the virtual environment. The device sends this information to the server, which then generates a corresponding dialogue scenario.
[0733] 3. Execute the dialogue:
[0734] For example, in response to a question about the fermentation process, "Tell me how alcohol is produced during fermentation," the server responds, "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide."
[0735] Specific examples
[0736] Consider a case where a user is learning about the beer fermentation process.
[0737] 1. Initial Setup:
[0738] The user launches the app, selects Japanese, and enters basic information. The device sends the information to the server, which stores it in a database and generates a learning plan optimized for the user.
[0739] 2. Process Selection:
[0740] The user selects "fermentation process" and the terminal sends this information to the server, which then generates a corresponding dialogue scenario and sends it to the terminal.
[0741] 3. Start a conversation:
[0742] The device starts the dialogue scenario by asking, "Do you know how alcohol is produced during fermentation?" The user responds, "I don't know. Please tell me." The device sends this to the server, which replies, "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide," which the device then relays.
[0743] 4. Further dialogue:
[0744] The user then asks a further question (e.g., "What kinds of yeast are there?"). The device recognizes the question and sends it to the server, which generates an appropriate response, which the device then relays to the user via voice.
[0745] Prompt Sentence Examples
[0746] Below is an example of a prompt sentence to input to the generative AI model.
[0747] User Question: How is alcohol produced during fermentation?
[0748] Correct Response: During the fermentation process, yeast converts sugars into alcohol and carbon dioxide.
[0749] This system allows users to deepen their practical language learning and intercultural understanding through interactive dialogue in a virtual environment using the beer brewing process as a subject.
[0750] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0751] Step 1:
[0752] Entering and receiving user information
[0753] The user launches the application and inputs the language they want to learn and basic information through smart glasses or a head-mounted display (HMD). The device receives this information and sends it to the server. The input data is the user's language of choice and personal information, and the output is the user's basic information and learning plan, which are stored in a database. The server receives this data and stores it in a database.
[0754] Step 2:
[0755] Dialogue scenario generation
[0756] The server generates a dialogue scenario based on the user's selected language and information about the beer brewing process. It uses a generative AI model such as ChatGPT to create a scenario for natural dialogue. The input data are details of the beer brewing process and the user's selected language, and the output is the generated dialogue scenario. The server generates the dialogue scenario and sends it to the device.
[0757] Step 3:
[0758] Building a virtual environment
[0759] Based on the received dialogue scenario, the device creates a virtual environment in the smart glasses or head-mounted display (HMD) worn by the user. The input data is the dialogue scenario and virtual environment setting information sent from the server, and the output is a visual virtual environment provided to the user. The device displays each stage of the beer brewing process in a virtual space.
[0760] Step 4:
[0761] Starting a conversation
[0762] The user selects a specific stage of the beer brewing process in the virtual environment and begins a dialogue. The device sends this selection information to the server, which then prepares a corresponding dialogue scenario. The input data is the information about the stage selected by the user, and the output is the dialogue scenario prepared by the server. The device plays the scenario aloud and presents questions to the user.
[0763] Step 5:
[0764] Speech Recognition and Processing
[0765] The user responds to the question by voice. The device converts the user's voice into text data using voice recognition software and sends the text data to the server. The input data is the user's voice data, and the output is text data. The server receives this data and uses it in the next step.
[0766] Step 6:
[0767] Response Generation
[0768] The server uses a generative AI model to generate an appropriate response based on the received text data. The input data is the text data received from the user, and the output is the generated response text. The server generates this response and sends it to the device.
[0769] Step 7:
[0770] Response output
[0771] The terminal converts the response text received from the server into speech using speech synthesis software and presents the response to the user. The input data is the text data received from the server, and the output is a voice message provided to the user. As a specific example, if the user asks, "Please tell me how alcohol is produced during the fermentation process," the terminal will respond by voice, "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide."
[0772] Step 8:
[0773] Record your learning progress
[0774] The device records the user's learning progress and sends it to the server, which stores this information in a database and updates the user's optimized learning plan. The input data is the user's learning progress information, and the output is the updated learning information stored in the database.
[0775] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0776] This invention is a system that aims to help users deepen their language learning and intercultural understanding through the beer brewing process, and by incorporating an emotion engine that recognizes the user's emotions, it provides a more personalized learning experience. The system involves three parties: a server, a terminal, and a user, and clarifies the roles of each part.
[0777] Server Processing
[0778] The server serves as the core of this system and provides the following specific functions:
[0779] 1. Receiving user information:
[0780] The server receives language selection information and basic information input by the user through the terminal.
[0781] 2. Generating dialogue scenarios:
[0782] The server generates a dialogue scenario based on the beer brewing process. This dialogue scenario is structured appropriately based on the language selected by the user. Advanced generative AI such as ChatGPT is used to achieve natural dialogue.
[0783] 3. Response Generation:
[0784] The server receives questions and responses from users and uses AI technology to generate appropriate responses.
[0785] 4. Emotional Information Processing:
[0786] The server receives the user's emotional information sent from the terminal and adjusts the content and tone of the response to the user based on this information.
[0787] 5. Database Update:
[0788] The server periodically adds the latest information about the beer production process and new dialogue scenarios to the database and provides them to users.
[0789] Terminal handling
[0790] The terminal provides the following specific functions as an interface with the user:
[0791] 1. User Interface:
[0792] The terminal acts as a virtual tour guide, displaying each stage of the beer production process in the user's language of choice.
[0793] 2. Speech recognition and output:
[0794] The terminal converts the user's voice input into text and also synthesizes the text received from the server into speech for the user to hear.
[0795] 3. Emotion recognition:
[0796] The device uses voice input and facial recognition technology to recognize the user's emotions and sends that information to a server.
[0797] 4. Data synchronization:
[0798] The device periodically synchronizes learning progress with the server, providing the user with the most up-to-date information.
[0799] User Behavior
[0800] The user performs the following specific actions through this system:
[0801] 1. Initial Settings and Selections:
[0802] Launch the app, select the language you want to learn, and enter your basic information.
[0803] 2. Selection of beer production process:
[0804] Within the app, users select the stage of the beer-making process they want to learn about (e.g., fermentation).
[0805] 3. Execute the dialogue:
[0806] The user interacts with the system at each stage of the process to learn. For example, in response to a question about the fermentation process, "Tell me how alcohol is produced during fermentation," the server responds, "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide."
[0807] Specific examples
[0808] Example: A user is learning about the beer fermentation process.
[0809] 1. Initial Setup:
[0810] The user launches the app, selects Japanese, and enters basic information. The device sends the information to the server, which stores it in a database and generates a learning plan optimized for the user.
[0811] 2. Process Selection:
[0812] The user selects "fermentation process" and the terminal sends this information to the server, which then generates a corresponding dialogue scenario and sends it to the terminal.
[0813] 3. Start a conversation:
[0814] The device starts the dialogue scenario by asking, "Do you know how alcohol is produced during the fermentation process?" The user responds, "I don't know. Please tell me." The device sends this to the server, which responds, "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide," which the device then relays in voice.
[0815] 4. Emotion Recognition and Processing:
[0816] If the user shows an emotional reaction to a question (e.g., a questioning or surprised expression), the device recognizes that emotion and sends it to the server. The server then adjusts the tone and content of the response based on the emotional information, and generates an appropriate response and sends it to the device.
[0817] 5. Continuing the dialogue:
[0818] The user asks a further question (e.g., "What types of yeast are there?"). The device recognizes the question and sends it to the server, which generates an appropriate response, which the device then speaks to the user.
[0819] In this way, users can deepen their language learning and intercultural understanding through natural dialogue while learning about the beer brewing process. The introduction of an emotion engine provides responses that correspond to the user's emotions, creating a more personalized learning experience.
[0820] The processing flow will be explained below.
[0821] Step 1:
[0822] The user launches the app. The device displays a startup screen, prompting the user to select a language and enter basic information.
[0823] Step 2:
[0824] The user inputs the desired language and basic information, which the terminal then sends to the server.
[0825] Step 3:
[0826] The server stores the received user information in a database, generates an optimal learning plan for the user, and sends it to the device.
[0827] Step 4:
[0828] Based on the learning plan received from the server, the device displays a selection screen for the beer brewing process, prompting the user to select the process they wish to learn.
[0829] Step 5:
[0830] The user selects "Fermentation process." The terminal transmits the user's selection to the server.
[0831] Step 6:
[0832] The server generates a dialogue scenario corresponding to the "fermentation process." It uses generative AI such as ChatGPT to build the dialogue scenario in the language selected by the user.
[0833] Step 7:
[0834] The server transmits the generated dialogue scenario to the terminal.
[0835] Step 8:
[0836] The device displays the dialogue scenario to the user and also starts voice output, asking the question, "Do you know how alcohol is produced during the fermentation process?"
[0837] Step 9:
[0838] The user responds to the question (e.g., "I don't understand. Please tell me."). The device recognizes the user's voice, converts it into text, and sends it to the server.
[0839] Step 10:
[0840] The server generates an appropriate response based on the user's response, using AI techniques to create answers in natural language (e.g., "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide.").
[0841] Step 11:
[0842] The server sends the generated response to the terminal.
[0843] Step 12:
[0844] The terminal synthesizes the response from the server into voice and plays it back to the user.
[0845] Step 13:
[0846] The user asks a follow-up question (e.g., "What kinds of yeast are there?"). The device uses voice recognition to convert the question into text and sends it to the server.
[0847] Step 14:
[0848] The server generates a response to the new question (e.g., "There are many different types of yeast, including ale yeast and lager yeast.") and sends it to the terminal.
[0849] Step 15:
[0850] The device synthesizes the response into speech and plays it back to the user.
[0851] Step 16:
[0852] The emotion engine recognizes the tone of voice and facial expressions used during the user's questions and responses, and the device analyzes this and transmits the user's emotional information to the server.
[0853] Step 17:
[0854] Based on the received emotional information, the server adjusts the next response to match the tone and content of the user's feelings.
[0855] Step 18:
[0856] The device will provide users with responses tailored to their emotions, enabling more personalized interactions.
[0857] Step 19:
[0858] The device periodically sends the user's learning progress to the server, which records the progress and uses it to optimize the learning content for the next session.
[0859] Step 20:
[0860] The server periodically updates the database, adding new dialogue scenarios and the latest information on beer production, and the terminal provides the new information to the user.
[0861] Example 2
[0862] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0863] Traditional language learning systems struggled to provide a personalized experience based on users' interests and emotions. They also lacked dynamic learning content that incorporated the latest information. As a result, it was difficult to maintain learners' motivation, and effective learning could not be expected.
[0864] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0865] In this invention, the server includes means for receiving language selection information and basic information input by a user, means for generating a dialogue scenario related to the beer brewing process, means for providing the generated dialogue scenario in the user's language of choice, means for recognizing questions and responses input by the user and generating appropriate replies, means for analyzing the user's emotional state and adjusting the content and tone of the replies, means for outputting the generated replies to the user, means for adding the latest information to the database, and means for periodically updating the information in the database and providing it to the user. This provides a personalized learning experience based on the user's interests and emotions, and enables dynamic learning incorporating the latest information, thereby enabling effective language learning.
[0866] The "means for receiving language selection information and basic information input by the user" refers to an interface for receiving the language selected by the user through the application and basic information such as name, language level, and learning purpose.
[0867] The "means for generating dialogue scenarios related to the beer production process" is a system that uses a generative AI model or the like to generate dialogue-style scenarios based on each stage of beer production.
[0868] The "means for providing the generated dialogue scenario in a language selected by the user" refers to a means for displaying or audibly outputting the generated dialogue scenario in a language selected by the user.
[0869] "Means for recognizing questions and responses entered by users and generating appropriate responses" refers to a system that utilizes AI technology to recognize questions and responses from users in text or voice and generate responses to them.
[0870] "Means for analyzing the user's emotional state and adjusting the content and tone of the response" refers to means for detecting the user's emotional state using voice analysis or facial recognition technology and adjusting the content and tone of the response according to that emotion.
[0871] The "means for outputting the generated response to the user" refers to a means for providing the generated response from the server to the user in the form of voice or text.
[0872] The "means for adding the latest information to the database" is a means for periodically collecting the latest information on the beer production process and adding it to the database.
[0873] "Means for periodically updating the information in the database and providing it to users" refers to means for keeping the information in the database up to date and providing that information to users.
[0874] This invention is a system that aims to help users deepen their language learning and intercultural understanding through the beer brewing process, and by incorporating an emotion engine that recognizes the user's emotions, it provides a more personalized learning experience. The system involves three parties: a server, a terminal, and a user, and clarifies the roles of each part.
[0875] Specific server processing
[0876] The server serves as the core of this system and provides the following specific functions:
[0877] 1. Receiving user information:
[0878] The server receives language selection information and basic information entered by the user through the terminal, which is stored in a database and used to create a user profile.
[0879] 2. Generating dialogue scenarios:
[0880] The server generates a dialogue scenario based on the beer brewing process. This dialogue scenario is structured appropriately based on the language selected by the user. Natural dialogue is achieved using generative AI models such as ChatGPT.
[0881] 3. Response Generation:
[0882] The server receives questions and responses from users and uses AI technology to generate appropriate responses, allowing users to learn specialized knowledge about beer brewing.
[0883] 4. Emotional Information Processing:
[0884] The server receives the user's emotional information sent from the terminal and adjusts the content and tone of the response to the user based on this information.
[0885] 5. Database Update:
[0886] The server periodically adds the latest information about the beer production process and new dialogue scenarios to the database and provides them to users.
[0887] Specific processing of the terminal
[0888] The terminal provides the following specific functions as an interface with the user:
[0889] 1. User Interface:
[0890] The terminal acts as a virtual tour guide, displaying each stage of the beer production process in the user's language of choice.
[0891] 2. Speech recognition and output:
[0892] The terminal converts the user's voice input into text and also synthesizes the text received from the server into speech for the user to hear.
[0893] 3. Emotion recognition:
[0894] The device uses voice input and facial recognition technology to recognize the user's emotions and sends that information to a server.
[0895] 4. Data synchronization:
[0896] The device periodically synchronizes learning progress with the server, providing the user with the most up-to-date information.
[0897] Specific user actions
[0898] The user will perform the following specific actions through this system:
[0899] 1. Initial Settings and Selections:
[0900] Users launch the app, select the language they want to learn, and enter basic information. The device sends this information to the server, which stores it in a database and generates a personalized learning plan.
[0901] 2. Selection of beer production process:
[0902] Within the app, users select the stage of the beer brewing process they want to learn about (e.g., the fermentation stage). The device sends this information to the server, which then generates a corresponding dialogue scenario and sends it to the device.
[0903] 3. Execute the dialogue:
[0904] The user interacts with the system at each stage of the process. For example, in response to a question about the fermentation process, "Tell me how alcohol is produced during fermentation," the server responds with, "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide." For example, the user can ask a question using the following prompt:
[0905] "Please explain the fermentation process of beer in detail."
[0906] What types of yeast are there?
[0907] summary
[0908] The system is designed to enable users to deepen their language learning and intercultural understanding through natural dialogue while learning about the beer brewing process. The system incorporates an emotion engine to provide responses based on the user's emotions, creating a more personalized learning experience. The system also adds the latest information to the database and provides it to users, ensuring they are constantly learning new knowledge.
[0909] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0910] Program processing flow
[0911] Step 1: Enter and submit user information
[0912] Specific behavior:
[0913] 1. Input: The user launches the app, selects Japanese or another language they want to learn, and enters basic information such as their name, language level, and learning purpose.
[0914] 2. Data processing: The device collects this information and converts it into the appropriate format.
[0915] 3. Send: The device creates and sends an API request to send the formatted information to the server.
[0916] 4. Output: The server saves the user's basic information and language preference in a database and creates a user profile.
[0917] Step 2: Generate dialogue scenarios
[0918] Specific behavior:
[0919] 1. Input: The server retrieves information about the beer production process from a database.
[0920] 2. Data processing: Based on the information obtained, the server converts the prompt text into a format suitable for sending to the generative AI model.
[0921] 3. Generation: A generative AI model like ChatGPT generates a dialogue scenario based on the prompt.
[0922] 4. Output: The server receives the generated dialogue scenario, converts it into an appropriate format, and sends it to the terminal to be presented in the language selected by the user.
[0923] Step 3: Handling User Questions and Responses
[0924] Specific behavior:
[0925] 1. Input: The user enters questions and responses by voice or text at each stage of the process, for example, "Tell me how alcohol is produced during fermentation."
[0926] 2. Data processing: The device uses voice recognition technology to convert the voice into text and sends the text data to the server.
[0927] 3. Generation: The server analyzes the user's question, sends it as a prompt to the generative AI model, and generates an appropriate response.
[0928] 4. Output: The server receives the generated response and sends it to the terminal, which then provides the response to the user using voice synthesis technology.
[0929] Step 4: Recognizing and processing emotional information
[0930] Specific behavior:
[0931] 1. Input: The user asks questions and responds to the device, and emotions are input from facial expressions and tone of voice.
[0932] 2. Data processing: The device uses voice analysis and facial recognition technology to analyze the user's emotional state, for example, detecting expressions of doubt or surprise.
[0933] 3. Send: The device sends the emotion analysis results to the server.
[0934] 4. Processing: The server adjusts the tone and content of the response based on the emotional information, generates an appropriate response, and sends it to the device.
[0935] 5. Output: The device communicates the adjusted response received from the server to the user via voice synthesis.
[0936] Step 5: Update the database
[0937] Specific behavior:
[0938] 1. Input: The server retrieves up-to-date information about the beer production process from an external resource.
[0939] 2. Data processing: The server converts the acquired information into a format for adding to the database.
[0940] 3. Add: The server adds new information to the database and updates existing information.
[0941] 4. Output: Based on the updated information, the server generates a new dialogue scenario and sends it to the terminal to provide the user with up-to-date information.
[0942] (Application example 2)
[0943] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0944] The purpose of this invention is to provide a language learning system based on the beer brewing process that can personalize the user's learning experience and provide an interactive and effective learning environment. In particular, the invention aims to deepen the user's interest and understanding in learning by recognizing the user's emotions and adjusting the content and tone of responses accordingly, and to provide an experience based on actual procedures and culture by visually displaying the beer brewing process in a virtual environment.
[0945] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving language selection information and basic information input by the user, means for generating a dialogue scenario related to the beer brewing process, means for providing the generated dialogue scenario in the user's language, means for recognizing questions and responses input by the user and generating appropriate replies, means for outputting the generated replies to the user, means for recognizing emotions and adjusting the content and tone of the responses, and means for visually displaying the beer brewing process in a virtual environment. This enables users to deepen their language learning and intercultural understanding through a natural dialogue format through the beer brewing process. Furthermore, the use of emotion recognition and a virtual environment makes the learning experience more interactive and personalized.
[0946] "Means for receiving user-entered language selection information and basic information" refers to a device for entering the user's preferred language and typical profile information (e.g., name, age, affiliation, etc.) and the function for transmitting that information to the system.
[0947] "Means for generating dialogue scenarios related to the beer production process" refers to a function for generating scenarios for dialogue in a language selected by the learner based on each stage of beer production.
[0948] "Means for providing the generated dialogue scenario in a language selected by the user" refers to an interface and device for translating the generated dialogue scenario into a language selected by the user and effectively providing it.
[0949] "Means of recognizing questions and responses entered by the user and generating appropriate responses" refers to a function that uses voice recognition technology and natural language processing technology to understand the content entered by the user and automatically generate appropriate responses based on that.
[0950] "Means for outputting the generated response to the user" refers to an output device or software that conveys the system-generated response to the user in audio or text form.
[0951] "Means for recognizing emotions and adjusting the content and tone of responses" refers to a function that analyzes emotions from the user's facial expressions and voice, and optimizes the system's response content and speaking style according to those emotions.
[0952] "Means for visually displaying the beer production process in a virtual environment" refers to a function that uses virtual reality or augmented reality technology to display the production process so that the user can visually understand each stage of beer production.
[0953] This invention is a system that aims to help users deepen their language learning and intercultural understanding through the beer brewing process, and by incorporating an emotion engine that recognizes the user's emotions, it provides a more personalized learning experience. Here, we will explain in detail how to specifically implement the invention.
[0954] System Overview
[0955] This system is realized by the collaboration of three parties: the server, the terminal, and the user, each fulfilling their respective roles. The server is responsible for all main processing, while the terminal functions as an interface with the user, allowing the user to learn through dialogue.
[0956] Program processing
[0957] Server Processing
[0958] 1. Receiving user information:
[0959] The server receives the language selection information and basic information entered by the user through the terminal, and this function is used to securely store the entered data and generate the necessary study plan.
[0960] 2. Generating dialogue scenarios:
[0961] The server generates a dialogue scenario based on the beer brewing process. This dialogue scenario is structured appropriately based on the language selected by the user. It uses OpenAI's generative AI model to achieve natural dialogue.
[0962] 3. Response Generation:
[0963] The server receives questions and responses from users and uses natural language processing technology to generate appropriate responses, using the latest generative AI models.
[0964] 4. Emotional Information Processing:
[0965] The server receives the user's emotional information sent from the terminal and adjusts the content and tone of the response to the user based on this information.
[0966] 5. Virtual environment display:
[0967] The server generates data to visually display each stage of the beer production process and sends it to the device, allowing users to experience it in a realistic way through visual content.
[0968] Terminal handling
[0969] 1. User Interface:
[0970] The terminal acts as a virtual tour guide, displaying each stage of the beer production process in the user's language of choice, with an intuitive and easy-to-use interface.
[0971] 2. Speech recognition and output:
[0972] The device converts the user's voice input into text and synthesizes the text received from the server into speech for the user to hear. The libraries used are "speech_recognition" and "pyttsx3".
[0973] 3. Emotion recognition:
[0974] The device recognizes the user's emotions using voice input and facial recognition technology, and sends the information to the server. Emotion recognition is performed using OpenCV and a pre-trained emotion recognition model.
[0975] 4. Data synchronization:
[0976] The device periodically synchronizes learning progress with the server, providing the user with the most up-to-date information.
[0977] User Behavior
[0978] 1. Initial Settings and Selections:
[0979] Users launch the app, select the language they want to learn, and enter basic information, which the device then sends to the server.
[0980] 2. Selection of beer production process:
[0981] Within the app, users select the stage of the beer-making process they want to learn about (e.g., fermentation), and the device sends this information to the server, which then generates a corresponding dialogue scenario.
[0982] 3. Execute the dialogue:
[0983] The user interacts with the system at each stage of the process to learn. For example, in response to a question about the fermentation process, "Tell me how alcohol is produced during fermentation," the server responds, "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide."
[0984] 4. Emotion Recognition and Processing:
[0985] If the user shows an emotional reaction to a question (e.g., a questioning or surprised expression), the device recognizes that emotion and transmits it to the server. The server then adjusts the tone and content of the response based on the emotional information, and generates an appropriate response and transmits it to the device.
[0986] 5. Continuing the dialogue:
[0987] If the user asks a further question (e.g., "What types of yeast are there?"), the device recognizes the question and sends it to the server, which generates an appropriate response, which the device then speaks to the user.
[0988] Prompt Sentence Examples
[0989] "A Guide to the Beer Making Process: Tell Me About Fermentation."
[0990] "A Guide to the Beer Making Process: Explain the Malting Steps."
[0991] In this way, users can deepen their language learning and intercultural understanding through natural dialogue while learning about the beer brewing process. The introduction of an emotion engine provides responses that correspond to the user's emotions, creating a more personalized learning experience.
[0992] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0993] Step 1:
[0994] The server receives language selection information and basic information entered by the user through the device, including profile information such as the user's preferred language, name, and age. The received information is stored in a database and used to generate a learning plan.
[0995] Step 2:
[0996] The server generates a dialogue scenario about the beer brewing process. It uses a generative AI model to create a natural dialogue scenario based on the language selected by the user. This dialogue scenario contains content related to the stage of the process the user wants to learn about (e.g., fermentation). The input used here is the user information saved in step 1, and the output is the generated dialogue scenario.
[0997] Step 3:
[0998] The terminal provides the generated dialogue scenario in the language selected by the user. The user interface (UI) displays the dialogue scenario visually and audibly, allowing the user to begin the dialogue. The input here is the dialogue scenario sent from the server, and the output is the dialogue scenario displayed on the user interface.
[0999] Step 4:
[1000] The user inputs a question or response into the terminal. Speech recognition technology is used to convert the user's voice input into text, and the text is sent to the server. The input is the user's voice, and the output is the user's question or response converted into text.
[1001] Step 5:
[1002] The server recognizes the questions and responses entered by the user and generates an appropriate response. It uses a generative AI model to perform natural language processing and generate an appropriate response to the user's question. The input is the text received in step 4, and the output is the generated response.
[1003] Step 6:
[1004] The terminal outputs the response received from the server to the user as voice. It uses speech synthesis technology (pyttsx3) to convey information in a way that is easy for the user to understand. The input is the text of the response sent from the server, and the output is the voice response.
[1005] Step 7:
[1006] The device recognizes the user's emotions and sends the information to the server. Emotion recognition uses OpenCV and a pre-trained emotion recognition model. The input is the user's facial image and voice, and the output is the recognized emotion information.
[1007] Step 8:
[1008] The server adjusts the content and tone of the response based on the emotional information. It changes the tone and content of the response depending on the emotion expressed by the user, generating a more personalized response. The input is the emotional information received in step 7, and the output is the adjusted response.
[1009] Step 9:
[1010] The device outputs the adjusted response to the user as voice, again using speech synthesis technology to provide an adapted response to the user. The input is the adjusted response text sent from the server, and the output is the voice response.
[1011] Step 10:
[1012] The terminal visually displays the beer brewing process in a virtual environment. Using virtual reality and augmented reality technology, the user can visually experience the beer brewing process. The input is visual data sent from the server, and the output is the beer brewing process displayed in the virtual environment.
[1013] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1014] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1015] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1016] [Third embodiment]
[1017] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1018] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[1019] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1020] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1021] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1022] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1023] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1024] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1025] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1026] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1027] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1028] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1029] This invention relates to a system that allows users to deepen their language learning and intercultural understanding through the beer brewing process. The system involves three components: a server, a terminal, and a user, and clarifies the roles of each component.
[1030] Server Processing
[1031] The server serves as the core of this system and provides the following specific functions:
[1032] 1. Receiving user information:
[1033] The server receives language selection information and basic information input by the user through the terminal.
[1034] 2. Generating dialogue scenarios:
[1035] The server generates a dialogue scenario based on the beer brewing process. This dialogue scenario is structured appropriately based on the language selected by the user. Advanced generative AI such as ChatGPT is used to achieve natural dialogue.
[1036] 3. Response Generation:
[1037] The server receives questions and responses from users and uses AI technology to generate appropriate responses.
[1038] 4. Database Update:
[1039] The server periodically adds the latest information about the beer production process and new dialogue scenarios to the database and provides them to users.
[1040] Terminal handling
[1041] The terminal provides the following specific functions as an interface with the user:
[1042] 1. User Interface:
[1043] The terminal acts as a virtual tour guide, displaying each stage of the beer production process in the user's language of choice.
[1044] 2. Speech recognition and output:
[1045] The terminal converts the user's voice input into text and also synthesizes the text received from the server into speech for the user to hear.
[1046] 3. Data synchronization:
[1047] The device periodically synchronizes learning progress with the server, providing the user with the most up-to-date information.
[1048] User Behavior
[1049] The user performs the following specific actions through this system:
[1050] 1. Initial Settings and Selections:
[1051] Launch the app, select the language you want to learn, and enter your basic information.
[1052] 2. Selection of beer production process:
[1053] Within the app, users select the stage of the beer-making process they want to learn about (e.g., fermentation).
[1054] 3. Execute the dialogue:
[1055] The user interacts with the system at each stage of the process to learn. For example, in response to a question about the fermentation process, "Tell me how alcohol is produced during fermentation," the server responds, "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide."
[1056] Specific examples
[1057] Example: A user is learning about the beer fermentation process.
[1058] 1. Initial Setup:
[1059] The user launches the app, selects Japanese, and enters basic information. The device sends the information to the server, which stores it in a database and generates a learning plan optimized for the user.
[1060] 2. Process Selection:
[1061] The user selects "fermentation process" and the terminal sends this information to the server, which then generates a corresponding dialogue scenario and sends it to the terminal.
[1062] 3. Start a conversation:
[1063] The device starts the dialogue scenario by asking, "Do you know how alcohol is produced during the fermentation process?" The user responds, "I don't know. Please tell me." The device sends this to the server, which responds, "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide," which the device then relays in voice.
[1064] 4. Continuing the dialogue:
[1065] The user asks a further question (e.g., "What types of yeast are there?"). The device recognizes the question and sends it to the server, which generates an appropriate response, which the device then speaks to the user.
[1066] This system allows users to deepen their language learning and intercultural understanding through natural dialogue while using the beer brewing process as a subject, thereby increasing their motivation to learn.
[1067] The processing flow will be explained below.
[1068] Step 1:
[1069] The user launches the app. The device displays a startup screen, prompting the user to select a language and enter basic information.
[1070] Step 2:
[1071] The user inputs the desired language and basic information, which the terminal then sends to the server.
[1072] Step 3:
[1073] The server stores the received user information in a database, generates an optimal learning plan for the user, and sends it to the device.
[1074] Step 4:
[1075] Based on the learning plan received from the server, the device displays a selection screen for the beer brewing process, prompting the user to select the process they wish to learn.
[1076] Step 5:
[1077] The user selects "Fermentation process." The terminal transmits the user's selection to the server.
[1078] Step 6:
[1079] The server generates a dialogue scenario corresponding to the "fermentation process." It uses generative AI such as ChatGPT to build the dialogue scenario in the language selected by the user.
[1080] Step 7:
[1081] The server transmits the generated dialogue scenario to the terminal.
[1082] Step 8:
[1083] The device displays a dialogue scenario to the user and also starts voice output, asking questions such as, "Do you know how alcohol is produced during the fermentation process?"
[1084] Step 9:
[1085] The user responds to the question (e.g., "I don't understand. Please tell me."). The device recognizes the user's voice, converts it into text, and sends it to the server.
[1086] Step 10:
[1087] The server generates an appropriate response based on the user's response, using AI techniques to create answers in natural language (e.g., "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide.").
[1088] Step 11:
[1089] The server sends the generated response to the terminal.
[1090] Step 12:
[1091] The terminal synthesizes the response from the server into voice and plays it back to the user.
[1092] Step 13:
[1093] The user continues to ask questions (e.g., "What kinds of yeast are there?"). The device converts the question into text using voice recognition and sends it to the server.
[1094] Step 14:
[1095] The server generates a response to the new question (e.g., "There are many different types of yeast, including ale yeast and lager yeast.") and sends it to the terminal.
[1096] Step 15:
[1097] The device synthesizes the response into voice and plays it back to the user. This process is repeated to advance learning in an interactive format.
[1098] Step 16:
[1099] The device periodically sends the user's learning progress to the server, which records the progress and uses it to optimize the learning content for the next session.
[1100] Step 17:
[1101] The server periodically updates the database, adding new dialogue scenarios and the latest information on beer production, and the terminal provides the new information to the user.
[1102] Example 1
[1103] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1104] Conventional language learning systems struggle to sustain users' interest and attention, leading to a decline in motivation to learn. Furthermore, most language learning systems simply focus on language acquisition, failing to incorporate cultural background and practical knowledge, making it difficult to deepen intercultural understanding. Furthermore, responses to user questions are mechanical, making it difficult to achieve natural dialogue. This tends to reduce learning efficiency and user satisfaction.
[1105] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1106] In this invention, the server includes means for receiving language selection information and basic information input by the user, means for generating a dialogue scenario related to a manufacturing process such as a fermentation process, means for providing the generated dialogue scenario in the user's selected language, means for recognizing questions and responses input by the user and generating appropriate replies, and means for outputting the generated replies to the user. This makes it possible to deepen language learning and intercultural understanding in a natural dialogue format while maintaining the user's interest, thereby improving learning efficiency and user satisfaction.
[1107] "User" refers to an individual who uses this system to input information and interact in order to deepen language learning and intercultural understanding.
[1108] "Language selection information" is information that a user inputs to the system when selecting the language they wish to learn.
[1109] "Basic information" refers to personal information such as name, age, and interests that a user provides to the system.
[1110] The "server" is the core of this system, and refers to a computer system that processes information received from the user and generates dialogue scenarios and responses.
[1111] The "manufacturing process" specifically refers to each stage of production activities, such as the fermentation process of beer, and dialogue scenarios are generated based on this theme.
[1112] A "dialogue scenario" is a planned conversation content for progressing learning and dialogue, which is generated based on the manufacturing process selected by the user.
[1113] "Response" refers to the words or documents generated by the system in response to a user's question or response.
[1114] "Speech recognition" is a technology that converts a user's voice input into text data.
[1115] "Speech synthesis" is a technology that converts text data into speech and reads it out to the user.
[1116] "Data Server" refers to the computer system used to record and store users' learning progress and latest information.
[1117] "Database" refers to a data storage system for systematically storing collected information and generated scenarios.
[1118] This invention is a system that allows users to deepen their language learning and intercultural understanding through the manufacturing process. This system involves three components: a server, a terminal, and a user, and each component clearly divides its role.
[1119] Server Processing
[1120] The server serves as the core of this system and provides the following specific functions:
[1121] 1. Receiving user information:
[1122] The server receives the language selection information and basic information entered by the user through the terminal, and records the information in a user profile database using specific software such as an HTTP server (e.g., Nginx) and a database server (e.g., MySQL).
[1123] 2. Generating dialogue scenarios:
[1124] The server generates a dialogue scenario based on a manufacturing process (e.g., a fermentation process). Here, generative AI (e.g., ChatGPT) is used to create a natural and interactive dialogue scenario. For example, the following prompts are used:
[1125] The user wants to learn about the "fermentation process." Generate a clear explanation of the fermentation process.
[1126] 3. Response Generation:
[1127] The server receives questions and responses from the user and uses generative AI to generate appropriate responses. For example, if the user asks, "Tell me how alcohol is produced during fermentation," the server sends the following prompt to the generative AI:
[1128] User: How is alcohol produced during fermentation?
[1129] The generative AI replies, "During the fermentation process, yeast converts sugars into alcohol and carbon dioxide."
[1130] 4. Database Update:
[1131] The server periodically adds the latest information about the manufacturing process and new interaction scenarios to the database and provides them to users. The latest information is collected through web crawlers and external APIs and added to the database management system (DBMS).
[1132] Terminal handling
[1133] The terminal provides the following specific functions as an interface with the user:
[1134] 1. User Interface:
[1135] The device acts as a virtual tour guide, displaying each stage of the manufacturing process in the user's language of choice, and uses cross-platform frameworks such as React Native for its software.
[1136] 2. Speech recognition and output:
[1137] The device uses speech recognition technology (e.g., Google Cloud Speech-to-Text API) to convert the user's voice input into text, and also converts the text received from the server into speech using speech synthesis technology (e.g., Amazon Polly) and provides it to the user.
[1138] 3. Data synchronization:
[1139] The device periodically synchronizes its learning progress with the server to provide the user with the latest information, using a long-lived connection over WebSocket or HTTP.
[1140] User Behavior
[1141] The user performs the following specific actions through this system:
[1142] 1. Initial Settings and Selections:
[1143] Users launch the app, select the language they want to learn, and enter basic information. The device sends this information to the server, which stores it in a database. The server then generates a personalized learning plan for the user.
[1144] 2. Selection of beer production process:
[1145] Within the app, the user selects the stage of the manufacturing process they want to learn about (e.g., the fermentation stage). The device sends this information to the server, which then generates a corresponding dialogue scenario and sends it to the device.
[1146] 3. Execute the dialogue:
[1147] The user interacts with the system at each stage of the process to learn: for example, if a user asks, "Tell me how alcohol is produced during fermentation," the server responds, "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide."
[1148] 4. Continuing the dialogue:
[1149] If the user asks a further question (e.g., "What types of yeast are there?"), the device recognizes this question and sends it to the server, which generates an appropriate response, which the device then speaks to the user.
[1150] This system allows users to deepen their language learning and intercultural understanding through a natural dialogue format while using manufacturing processes as a subject, thereby increasing their motivation to learn.
[1151] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1152] Step 1: User launches the app
[1153] Input: A user taps on a device to launch an app.
[1154] Output: A language selection screen will appear.
[1155] What happens: When a user launches the app, the device loads the app's initial screen. The app is built with a cross-platform framework such as React Native and prompts the user to select a language.
[1156] Step 2: User enters language and basic information
[1157] Input: Enter your language preference and basic information (name, age, interests, etc.).
[1158] Output: The entered information is sent to the server and stored in a database.
[1159] Specific operation: The user enters the language and basic information in the input form and presses the "Submit" button. The terminal sends this as an HTTP POST request to the server, which receives the request via Nginx and stores it in the MySQL database.
[1160] Step 3: The server generates the lesson plan
[1161] Input: User's language preference and basic information.
[1162] Output: User-optimized learning plan data.
[1163] Specific operation: The server generates an optimal learning plan based on the received user information, referencing the data in the database. The generated learning plan is returned to the device in JSON format.
[1164] Step 4: User selects stage in manufacturing process
[1165] Input: The stage of the manufacturing process you want to learn about (e.g., fermentation process).
[1166] Output: Stage selection information is sent to the server.
[1167] Specific operation: When a user selects "Fermentation process" in the app, the device sends this information in JSON format to the server as an HTTP POST request. The server receives this information.
[1168] Step 5: The server generates the dialogue scenario
[1169] Input: Manufacturing process stage information.
[1170] Output: Dialogue scenario data.
[1171] Specific operation: The server sends prompts to generative AI models such as ChatGPT based on the manufacturing process. For example, it sends the following prompts:
[1172] The user wants to learn about the "fermentation process." Generate a clear explanation of the fermentation process.
[1173] The server receives the scenario generated by ChatGPT and returns it to the terminal in JSON format.
[1174] Step 6: The device presents the dialogue scenario
[1175] Input: Dialogue scenario data.
[1176] Output: Presents the dialogue scenario to the user in voice and text.
[1177] Specific operation: The device analyzes the received dialogue scenario and presents it to the user in both voice and text format. It converts the text into speech using speech synthesis technology (e.g., Amazon Polly).
[1178] Step 7: User enters question and response
[1179] Input: User voice or text input.
[1180] Output: The entered questions and responses are sent to the server in text format.
[1181] Specific operation: The user responds with something like "I don't understand. Please tell me." The device uses the Google Cloud Speech-to-Text API to convert the speech into text and send it to the server.
[1182] Step 8: The server generates a response
[1183] Input: The user's question and response.
[1184] Output: The generated response data.
[1185] How it works: The server receives the user's question and sends the following prompt to the generative AI (e.g. ChatGPT):
[1186] User: How is alcohol produced during fermentation?
[1187] The server receives the response generated by ChatGPT: "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide." and sends it to the terminal.
[1188] Step 9: The device presents a response
[1189] Input: The generated response data.
[1190] Output: Presents the response to the user in audio and text.
[1191] Specific operation: The device analyzes the response received from the server and uses speech synthesis technology to present it to the user in the form of voice and text.
[1192] This series of steps allows users to deepen their language learning and intercultural understanding through a natural, interactive way during the manufacturing process.
[1193] (Application example 1)
[1194] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1195] Conventional language learning systems have limited content that attracts users' interest, making it difficult to maintain motivation. In particular, they are limited to learning grammar and vocabulary, with few opportunities to learn cultural background or practical applications. Another problem is that the lack of interactive learning makes it difficult to improve practical language skills.
[1196] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1197] In this invention, the server includes means for receiving language selection information and basic information input by a user, means for generating a dialogue scenario related to the beer brewing process, means for providing the generated dialogue scenario in a language selected by the user, means for recognizing questions and responses input by the user and generating appropriate responses, means for outputting the generated responses to the user, means for visually representing the beer brewing process in a virtual environment, and means for the user to learn a language through dialogue in the virtual environment. This enables the user to deepen practical language learning and intercultural understanding through interactive dialogue in the virtual environment using the beer brewing process as a subject.
[1198] "Language selection information and basic information entered by the user" refers to the language used for learning that the user selects for the system and basic personal information.
[1199] The "means for generating a dialogue scenario" refers to a system component that has the function of automatically creating a dialogue scenario with the user based on information about the beer brewing process.
[1200] "Means for providing in a language selected by the user" refers to a system component that has the function of translating the generated dialogue scenario into a specific language selected by the user and presenting information in that language.
[1201] "Means for recognizing questions and responses and generating appropriate responses" refers to a system component that has the function of recognizing and analyzing questions and responses entered by users in voice or text format and automatically generating appropriate responses to them.
[1202] "Means for outputting the generated response to the user" refers to a system component that has the function of presenting the response generated by the system to the user in the form of voice or text.
[1203] "Means for visually representing the beer production process in a virtual environment" refers to a component of a system that uses virtual reality technology to visually recreate the beer production process and provide users with an interactive experience.
[1204] "Means for users to learn a language through interaction in a virtual environment" refers to a component of a system that has the function of allowing users to learn a language through interaction in an environment created using virtual reality technology.
[1205] This invention relates to a system that aims to deepen language learning and intercultural understanding through the beer brewing process. Specifically, it utilizes virtual reality technology and generative AI models to provide users with an interactive learning experience.
[1206] Server Processing
[1207] The server serves as the core of this system and provides the following specific functions:
[1208] 1. Receiving user information:
[1209] The server receives the language selection information and basic information input by the user through the terminal and stores them in a database.
[1210] 2. Dialogue scenario generation:
[1211] The server generates a dialogue scenario based on the beer brewing process, and the dialogue scenario is structured appropriately based on the language selected by the user. Advanced generative AI models such as ChatGPT are used to achieve natural dialogue.
[1212] 3. Response Generation:
[1213] The server receives questions and responses from users and uses generative AI technology to generate appropriate responses.
[1214] 4. Data synchronization:
[1215] The server records the user's learning progress and periodically adds the latest information and new dialogue scenarios to the database and provides them to the user.
[1216] Terminal handling
[1217] The terminal provides the following specific functions as an interface with the user:
[1218] 1. User Interface:
[1219] It acts as a virtual tour guide, showing each stage of the beer production process in the user's language of choice.
[1220] 2. Speech Recognition and Synthesis:
[1221] The system converts the user's voice input into text, synthesizes the text received from the server into speech, and plays it back to the user.
[1222] 3. Virtual Environment:
[1223] Users use smart glasses or a head-mounted display (HMD) to visually experience the beer production process in a virtual environment.
[1224] User Behavior
[1225] The user performs the following specific actions through this system:
[1226] 1. Initial Settings and Selections:
[1227] Launch the app, select the language you want to learn, and enter some basic information. The device then sends the information to the server, which stores it in a database.
[1228] 2. Selection of beer production process:
[1229] The user selects the stage of the beer production process (e.g., fermentation) they want to learn about in the virtual environment. The device sends this information to the server, which then generates a corresponding dialogue scenario.
[1230] 3. Execute the dialogue:
[1231] For example, in response to a question about the fermentation process, "Tell me how alcohol is produced during fermentation," the server responds, "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide."
[1232] Specific examples
[1233] Consider a case where a user is learning about the beer fermentation process.
[1234] 1. Initial Setup:
[1235] The user launches the app, selects Japanese, and enters basic information. The device sends the information to the server, which stores it in a database and generates a learning plan optimized for the user.
[1236] 2. Process Selection:
[1237] The user selects "fermentation process" and the terminal sends this information to the server, which then generates a corresponding dialogue scenario and sends it to the terminal.
[1238] 3. Start a conversation:
[1239] The device starts the dialogue scenario by asking, "Do you know how alcohol is produced during fermentation?" The user responds, "I don't know. Please tell me." The device sends this to the server, which replies, "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide," which the device then relays.
[1240] 4. Further dialogue:
[1241] The user then asks a further question (e.g., "What kinds of yeast are there?"). The device recognizes the question and sends it to the server, which generates an appropriate response, which the device then relays to the user via voice.
[1242] Prompt Sentence Examples
[1243] Below is an example of a prompt sentence to input to the generative AI model.
[1244] User Question: How is alcohol produced during fermentation?
[1245] Correct Response: During the fermentation process, yeast converts sugars into alcohol and carbon dioxide.
[1246] This system allows users to deepen their practical language learning and intercultural understanding through interactive dialogue in a virtual environment using the beer brewing process as a subject.
[1247] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1248] Step 1:
[1249] Entering and receiving user information
[1250] The user launches the application and inputs the language they want to learn and basic information through smart glasses or a head-mounted display (HMD). The device receives this information and sends it to the server. The input data is the user's language of choice and personal information, and the output is the user's basic information and learning plan, which are stored in a database. The server receives this data and stores it in a database.
[1251] Step 2:
[1252] Dialogue scenario generation
[1253] The server generates a dialogue scenario based on the user's selected language and information about the beer brewing process. It uses a generative AI model such as ChatGPT to create a scenario for natural dialogue. The input data are details of the beer brewing process and the user's selected language, and the output is the generated dialogue scenario. The server generates the dialogue scenario and sends it to the device.
[1254] Step 3:
[1255] Building a virtual environment
[1256] Based on the received dialogue scenario, the device creates a virtual environment in the smart glasses or head-mounted display (HMD) worn by the user. The input data is the dialogue scenario and virtual environment setting information sent from the server, and the output is a visual virtual environment provided to the user. The device displays each stage of the beer brewing process in a virtual space.
[1257] Step 4:
[1258] Starting a conversation
[1259] The user selects a specific stage of the beer brewing process in the virtual environment and begins a dialogue. The device sends this selection information to the server, which then prepares a corresponding dialogue scenario. The input data is the information about the stage selected by the user, and the output is the dialogue scenario prepared by the server. The device plays the scenario aloud and presents questions to the user.
[1260] Step 5:
[1261] Speech Recognition and Processing
[1262] The user responds to the question by voice. The device converts the user's voice into text data using voice recognition software and sends the text data to the server. The input data is the user's voice data, and the output is text data. The server receives this data and uses it in the next step.
[1263] Step 6:
[1264] Response Generation
[1265] The server uses a generative AI model to generate an appropriate response based on the received text data. The input data is the text data received from the user, and the output is the generated response text. The server generates this response and sends it to the device.
[1266] Step 7:
[1267] Response output
[1268] The terminal converts the response text received from the server into speech using speech synthesis software and presents the response to the user. The input data is the text data received from the server, and the output is a voice message provided to the user. As a specific example, if the user asks, "Please tell me how alcohol is produced during the fermentation process," the terminal will respond by voice, "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide."
[1269] Step 8:
[1270] Record your learning progress
[1271] The device records the user's learning progress and sends it to the server, which stores this information in a database and updates the user's optimized learning plan. The input data is the user's learning progress information, and the output is the updated learning information stored in the database.
[1272] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1273] This invention is a system that aims to help users deepen their language learning and intercultural understanding through the beer brewing process, and by incorporating an emotion engine that recognizes the user's emotions, it provides a more personalized learning experience. The system involves three parties: a server, a terminal, and a user, and clarifies the roles of each part.
[1274] Server Processing
[1275] The server serves as the core of this system and provides the following specific functions:
[1276] 1. Receiving user information:
[1277] The server receives language selection information and basic information input by the user through the terminal.
[1278] 2. Generating dialogue scenarios:
[1279] The server generates a dialogue scenario based on the beer brewing process. This dialogue scenario is structured appropriately based on the language selected by the user. Advanced generative AI such as ChatGPT is used to achieve natural dialogue.
[1280] 3. Response Generation:
[1281] The server receives questions and responses from users and uses AI technology to generate appropriate responses.
[1282] 4. Emotional Information Processing:
[1283] The server receives the user's emotional information sent from the terminal and adjusts the content and tone of the response to the user based on this information.
[1284] 5. Database Update:
[1285] The server periodically adds the latest information about the beer production process and new dialogue scenarios to the database and provides them to users.
[1286] Terminal handling
[1287] The terminal provides the following specific functions as an interface with the user:
[1288] 1. User Interface:
[1289] The terminal acts as a virtual tour guide, displaying each stage of the beer production process in the user's language of choice.
[1290] 2. Speech recognition and output:
[1291] The terminal converts the user's voice input into text and also synthesizes the text received from the server into speech for the user to hear.
[1292] 3. Emotion recognition:
[1293] The device uses voice input and facial recognition technology to recognize the user's emotions and sends that information to a server.
[1294] 4. Data synchronization:
[1295] The device periodically synchronizes learning progress with the server, providing the user with the most up-to-date information.
[1296] User Behavior
[1297] The user performs the following specific actions through this system:
[1298] 1. Initial Settings and Selections:
[1299] Launch the app, select the language you want to learn, and enter your basic information.
[1300] 2. Selection of beer production process:
[1301] Within the app, users select the stage of the beer-making process they want to learn about (e.g., fermentation).
[1302] 3. Execute the dialogue:
[1303] The user interacts with the system at each stage of the process to learn. For example, in response to a question about the fermentation process, "Tell me how alcohol is produced during fermentation," the server responds, "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide."
[1304] Specific examples
[1305] Example: A user is learning about the beer fermentation process.
[1306] 1. Initial Setup:
[1307] The user launches the app, selects Japanese, and enters basic information. The device sends the information to the server, which stores it in a database and generates a learning plan optimized for the user.
[1308] 2. Process Selection:
[1309] The user selects "fermentation process" and the terminal sends this information to the server, which then generates a corresponding dialogue scenario and sends it to the terminal.
[1310] 3. Start a conversation:
[1311] The device starts the dialogue scenario by asking, "Do you know how alcohol is produced during the fermentation process?" The user responds, "I don't know. Please tell me." The device sends this to the server, which responds, "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide," which the device then relays in voice.
[1312] 4. Emotion Recognition and Processing:
[1313] If the user shows an emotional reaction to a question (e.g., a questioning or surprised expression), the device recognizes that emotion and sends it to the server. The server then adjusts the tone and content of the response based on the emotional information, and generates an appropriate response and sends it to the device.
[1314] 5. Continuing the dialogue:
[1315] The user asks a further question (e.g., "What types of yeast are there?"). The device recognizes the question and sends it to the server, which generates an appropriate response, which the device then speaks to the user.
[1316] In this way, users can deepen their language learning and intercultural understanding through natural dialogue while learning about the beer brewing process. The introduction of an emotion engine provides responses that correspond to the user's emotions, creating a more personalized learning experience.
[1317] The processing flow will be explained below.
[1318] Step 1:
[1319] The user launches the app. The device displays a startup screen, prompting the user to select a language and enter basic information.
[1320] Step 2:
[1321] The user inputs the desired language and basic information, which the terminal then sends to the server.
[1322] Step 3:
[1323] The server stores the received user information in a database, generates an optimal learning plan for the user, and sends it to the device.
[1324] Step 4:
[1325] Based on the learning plan received from the server, the device displays a selection screen for the beer brewing process, prompting the user to select the process they wish to learn.
[1326] Step 5:
[1327] The user selects "Fermentation process." The terminal transmits the user's selection to the server.
[1328] Step 6:
[1329] The server generates a dialogue scenario corresponding to the "fermentation process." It uses generative AI such as ChatGPT to build the dialogue scenario in the language selected by the user.
[1330] Step 7:
[1331] The server transmits the generated dialogue scenario to the terminal.
[1332] Step 8:
[1333] The device displays the dialogue scenario to the user and also starts voice output, asking the question, "Do you know how alcohol is produced during the fermentation process?"
[1334] Step 9:
[1335] The user responds to the question (e.g., "I don't understand. Please tell me."). The device recognizes the user's voice, converts it into text, and sends it to the server.
[1336] Step 10:
[1337] The server generates an appropriate response based on the user's response, using AI techniques to create answers in natural language (e.g., "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide.").
[1338] Step 11:
[1339] The server sends the generated response to the terminal.
[1340] Step 12:
[1341] The terminal synthesizes the response from the server into voice and plays it back to the user.
[1342] Step 13:
[1343] The user asks a follow-up question (e.g., "What kinds of yeast are there?"). The device uses voice recognition to convert the question into text and sends it to the server.
[1344] Step 14:
[1345] The server generates a response to the new question (e.g., "There are many different types of yeast, including ale yeast and lager yeast.") and sends it to the terminal.
[1346] Step 15:
[1347] The device synthesizes the response into speech and plays it back to the user.
[1348] Step 16:
[1349] The emotion engine recognizes the tone of voice and facial expressions used during the user's questions and responses, and the device analyzes this and transmits the user's emotional information to the server.
[1350] Step 17:
[1351] Based on the received emotional information, the server adjusts the next response to match the tone and content of the user's feelings.
[1352] Step 18:
[1353] The device will provide users with responses tailored to their emotions, enabling more personalized interactions.
[1354] Step 19:
[1355] The device periodically sends the user's learning progress to the server, which records the progress and uses it to optimize the learning content for the next session.
[1356] Step 20:
[1357] The server periodically updates the database, adding new dialogue scenarios and the latest information on beer production, and the terminal provides the new information to the user.
[1358] Example 2
[1359] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1360] Traditional language learning systems struggled to provide a personalized experience based on users' interests and emotions. They also lacked dynamic learning content that incorporated the latest information. As a result, it was difficult to maintain learners' motivation, and effective learning could not be expected.
[1361] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1362] In this invention, the server includes means for receiving language selection information and basic information input by a user, means for generating a dialogue scenario related to the beer brewing process, means for providing the generated dialogue scenario in the user's language of choice, means for recognizing questions and responses input by the user and generating appropriate replies, means for analyzing the user's emotional state and adjusting the content and tone of the replies, means for outputting the generated replies to the user, means for adding the latest information to the database, and means for periodically updating the information in the database and providing it to the user. This provides a personalized learning experience based on the user's interests and emotions, and enables dynamic learning incorporating the latest information, thereby enabling effective language learning.
[1363] The "means for receiving language selection information and basic information input by the user" refers to an interface for receiving the language selected by the user through the application and basic information such as name, language level, and learning purpose.
[1364] The "means for generating dialogue scenarios related to the beer production process" is a system that uses a generative AI model or the like to generate dialogue-style scenarios based on each stage of beer production.
[1365] The "means for providing the generated dialogue scenario in a language selected by the user" refers to a means for displaying or audibly outputting the generated dialogue scenario in a language selected by the user.
[1366] "Means for recognizing questions and responses entered by users and generating appropriate responses" refers to a system that utilizes AI technology to recognize questions and responses from users in text or voice and generate responses to them.
[1367] "Means for analyzing the user's emotional state and adjusting the content and tone of the response" refers to means for detecting the user's emotional state using voice analysis or facial recognition technology and adjusting the content and tone of the response according to that emotion.
[1368] The "means for outputting the generated response to the user" refers to a means for providing the generated response from the server to the user in the form of voice or text.
[1369] The "means for adding the latest information to the database" is a means for periodically collecting the latest information on the beer production process and adding it to the database.
[1370] "Means for periodically updating the information in the database and providing it to users" refers to means for keeping the information in the database up to date and providing that information to users.
[1371] This invention is a system that aims to help users deepen their language learning and intercultural understanding through the beer brewing process, and by incorporating an emotion engine that recognizes the user's emotions, it provides a more personalized learning experience. The system involves three parties: a server, a terminal, and a user, and clarifies the roles of each part.
[1372] Specific server processing
[1373] The server serves as the core of this system and provides the following specific functions:
[1374] 1. Receiving user information:
[1375] The server receives language selection information and basic information entered by the user through the terminal, which is stored in a database and used to create a user profile.
[1376] 2. Generating dialogue scenarios:
[1377] The server generates a dialogue scenario based on the beer brewing process. This dialogue scenario is structured appropriately based on the language selected by the user. Natural dialogue is achieved using generative AI models such as ChatGPT.
[1378] 3. Response Generation:
[1379] The server receives questions and responses from users and uses AI technology to generate appropriate responses, allowing users to learn specialized knowledge about beer brewing.
[1380] 4. Emotional Information Processing:
[1381] The server receives the user's emotional information sent from the terminal and adjusts the content and tone of the response to the user based on this information.
[1382] 5. Database Update:
[1383] The server periodically adds the latest information about the beer production process and new dialogue scenarios to the database and provides them to users.
[1384] Specific processing of the terminal
[1385] The terminal provides the following specific functions as an interface with the user:
[1386] 1. User Interface:
[1387] The terminal acts as a virtual tour guide, displaying each stage of the beer production process in the user's language of choice.
[1388] 2. Speech recognition and output:
[1389] The terminal converts the user's voice input into text and also synthesizes the text received from the server into speech for the user to hear.
[1390] 3. Emotion recognition:
[1391] The device uses voice input and facial recognition technology to recognize the user's emotions and sends that information to a server.
[1392] 4. Data synchronization:
[1393] The device periodically synchronizes learning progress with the server, providing the user with the most up-to-date information.
[1394] Specific user actions
[1395] The user will perform the following specific actions through this system:
[1396] 1. Initial Settings and Selections:
[1397] Users launch the app, select the language they want to learn, and enter basic information. The device sends this information to the server, which stores it in a database and generates a personalized learning plan.
[1398] 2. Selection of beer production process:
[1399] Within the app, users select the stage of the beer brewing process they want to learn about (e.g., the fermentation stage). The device sends this information to the server, which then generates a corresponding dialogue scenario and sends it to the device.
[1400] 3. Execute the dialogue:
[1401] The user interacts with the system at each stage of the process. For example, in response to a question about the fermentation process, "Tell me how alcohol is produced during fermentation," the server responds with, "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide." For example, the user can ask a question using the following prompt:
[1402] "Please explain the fermentation process of beer in detail."
[1403] What types of yeast are there?
[1404] summary
[1405] The system is designed to enable users to deepen their language learning and intercultural understanding through natural dialogue while learning about the beer brewing process. The system incorporates an emotion engine to provide responses based on the user's emotions, creating a more personalized learning experience. The system also adds the latest information to the database and provides it to users, ensuring they are constantly learning new knowledge.
[1406] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1407] Program processing flow
[1408] Step 1: Enter and submit user information
[1409] Specific behavior:
[1410] 1. Input: The user launches the app, selects Japanese or another language they want to learn, and enters basic information such as their name, language level, and learning purpose.
[1411] 2. Data processing: The device collects this information and converts it into the appropriate format.
[1412] 3. Send: The device creates and sends an API request to send the formatted information to the server.
[1413] 4. Output: The server saves the user's basic information and language preference in a database and creates a user profile.
[1414] Step 2: Generate dialogue scenarios
[1415] Specific behavior:
[1416] 1. Input: The server retrieves information about the beer production process from a database.
[1417] 2. Data processing: Based on the information obtained, the server converts the prompt text into a format suitable for sending to the generative AI model.
[1418] 3. Generation: A generative AI model like ChatGPT generates a dialogue scenario based on the prompt.
[1419] 4. Output: The server receives the generated dialogue scenario, converts it into an appropriate format, and sends it to the terminal to be presented in the language selected by the user.
[1420] Step 3: Handling User Questions and Responses
[1421] Specific behavior:
[1422] 1. Input: The user enters questions and responses by voice or text at each stage of the process, for example, "Tell me how alcohol is produced during fermentation."
[1423] 2. Data processing: The device uses voice recognition technology to convert the voice into text and sends the text data to the server.
[1424] 3. Generation: The server analyzes the user's question, sends it as a prompt to the generative AI model, and generates an appropriate response.
[1425] 4. Output: The server receives the generated response and sends it to the terminal, which then provides the response to the user using voice synthesis technology.
[1426] Step 4: Recognizing and processing emotional information
[1427] Specific behavior:
[1428] 1. Input: The user asks questions and responds to the device, and emotions are input from facial expressions and tone of voice.
[1429] 2. Data processing: The device uses voice analysis and facial recognition technology to analyze the user's emotional state, for example, detecting expressions of doubt or surprise.
[1430] 3. Send: The device sends the emotion analysis results to the server.
[1431] 4. Processing: The server adjusts the tone and content of the response based on the emotional information, generates an appropriate response, and sends it to the device.
[1432] 5. Output: The device communicates the adjusted response received from the server to the user via voice synthesis.
[1433] Step 5: Update the database
[1434] Specific behavior:
[1435] 1. Input: The server retrieves up-to-date information about the beer production process from an external resource.
[1436] 2. Data processing: The server converts the acquired information into a format for adding to the database.
[1437] 3. Add: The server adds new information to the database and updates existing information.
[1438] 4. Output: Based on the updated information, the server generates a new dialogue scenario and sends it to the terminal to provide the user with up-to-date information.
[1439] (Application example 2)
[1440] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1441] The purpose of this invention is to provide a language learning system based on the beer brewing process that can personalize the user's learning experience and provide an interactive and effective learning environment. In particular, the invention aims to deepen the user's interest and understanding in learning by recognizing the user's emotions and adjusting the content and tone of responses accordingly, and to provide an experience based on actual procedures and culture by visually displaying the beer brewing process in a virtual environment.
[1442] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving language selection information and basic information input by the user, means for generating a dialogue scenario related to the beer brewing process, means for providing the generated dialogue scenario in the user's language, means for recognizing questions and responses input by the user and generating appropriate replies, means for outputting the generated replies to the user, means for recognizing emotions and adjusting the content and tone of the responses, and means for visually displaying the beer brewing process in a virtual environment. This enables users to deepen their language learning and intercultural understanding through a natural dialogue format through the beer brewing process. Furthermore, the use of emotion recognition and a virtual environment makes the learning experience more interactive and personalized.
[1443] "Means for receiving user-entered language selection information and basic information" refers to a device for entering the user's preferred language and typical profile information (e.g., name, age, affiliation, etc.) and the function for transmitting that information to the system.
[1444] "Means for generating dialogue scenarios related to the beer production process" refers to a function for generating scenarios for dialogue in a language selected by the learner based on each stage of beer production.
[1445] "Means for providing the generated dialogue scenario in a language selected by the user" refers to an interface and device for translating the generated dialogue scenario into a language selected by the user and effectively providing it.
[1446] "Means of recognizing questions and responses entered by the user and generating appropriate responses" refers to a function that uses voice recognition technology and natural language processing technology to understand the content entered by the user and automatically generate appropriate responses based on that.
[1447] "Means for outputting the generated response to the user" refers to an output device or software that conveys the system-generated response to the user in audio or text form.
[1448] "Means for recognizing emotions and adjusting the content and tone of responses" refers to a function that analyzes emotions from the user's facial expressions and voice, and optimizes the system's response content and speaking style according to those emotions.
[1449] "Means for visually displaying the beer production process in a virtual environment" refers to a function that uses virtual reality or augmented reality technology to display the production process so that the user can visually understand each stage of beer production.
[1450] This invention is a system that aims to help users deepen their language learning and intercultural understanding through the beer brewing process, and by incorporating an emotion engine that recognizes the user's emotions, it provides a more personalized learning experience. Here, we will explain in detail how to specifically implement the invention.
[1451] System Overview
[1452] This system is realized by the collaboration of three parties: the server, the terminal, and the user, each fulfilling their respective roles. The server is responsible for all main processing, while the terminal functions as an interface with the user, allowing the user to learn through dialogue.
[1453] Program processing
[1454] Server Processing
[1455] 1. Receiving user information:
[1456] The server receives the language selection information and basic information entered by the user through the terminal, and this function is used to securely store the entered data and generate the necessary study plan.
[1457] 2. Generating dialogue scenarios:
[1458] The server generates a dialogue scenario based on the beer brewing process. This dialogue scenario is structured appropriately based on the language selected by the user. It uses OpenAI's generative AI model to achieve natural dialogue.
[1459] 3. Response Generation:
[1460] The server receives questions and responses from users and uses natural language processing technology to generate appropriate responses, using the latest generative AI models.
[1461] 4. Emotional Information Processing:
[1462] The server receives the user's emotional information sent from the terminal and adjusts the content and tone of the response to the user based on this information.
[1463] 5. Virtual environment display:
[1464] The server generates data to visually display each stage of the beer production process and sends it to the device, allowing users to experience it in a realistic way through visual content.
[1465] Terminal handling
[1466] 1. User Interface:
[1467] The terminal acts as a virtual tour guide, displaying each stage of the beer production process in the user's language of choice, with an intuitive and easy-to-use interface.
[1468] 2. Speech recognition and output:
[1469] The device converts the user's voice input into text and synthesizes the text received from the server into speech for the user to hear. The libraries used are "speech_recognition" and "pyttsx3".
[1470] 3. Emotion recognition:
[1471] The device recognizes the user's emotions using voice input and facial recognition technology, and sends the information to the server. Emotion recognition is performed using OpenCV and a pre-trained emotion recognition model.
[1472] 4. Data synchronization:
[1473] The device periodically synchronizes learning progress with the server, providing the user with the most up-to-date information.
[1474] User Behavior
[1475] 1. Initial Settings and Selections:
[1476] Users launch the app, select the language they want to learn, and enter basic information, which the device then sends to the server.
[1477] 2. Selection of beer production process:
[1478] Within the app, users select the stage of the beer-making process they want to learn about (e.g., fermentation), and the device sends this information to the server, which then generates a corresponding dialogue scenario.
[1479] 3. Execute the dialogue:
[1480] The user interacts with the system at each stage of the process to learn. For example, in response to a question about the fermentation process, "Tell me how alcohol is produced during fermentation," the server responds, "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide."
[1481] 4. Emotion Recognition and Processing:
[1482] If the user shows an emotional reaction to a question (e.g., a questioning or surprised expression), the device recognizes that emotion and transmits it to the server. The server then adjusts the tone and content of the response based on the emotional information, and generates an appropriate response and transmits it to the device.
[1483] 5. Continuing the dialogue:
[1484] If the user asks a further question (e.g., "What types of yeast are there?"), the device recognizes the question and sends it to the server, which generates an appropriate response, which the device then speaks to the user.
[1485] Prompt Sentence Examples
[1486] "A Guide to the Beer Making Process: Tell Me About Fermentation."
[1487] "A Guide to the Beer Making Process: Explain the Malting Steps."
[1488] In this way, users can deepen their language learning and intercultural understanding through natural dialogue while learning about the beer brewing process. The introduction of an emotion engine provides responses that correspond to the user's emotions, creating a more personalized learning experience.
[1489] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1490] Step 1:
[1491] The server receives language selection information and basic information entered by the user through the device, including profile information such as the user's preferred language, name, and age. The received information is stored in a database and used to generate a learning plan.
[1492] Step 2:
[1493] The server generates a dialogue scenario about the beer brewing process. It uses a generative AI model to create a natural dialogue scenario based on the language selected by the user. This dialogue scenario contains content related to the stage of the process the user wants to learn about (e.g., fermentation). The input used here is the user information saved in step 1, and the output is the generated dialogue scenario.
[1494] Step 3:
[1495] The terminal provides the generated dialogue scenario in the language selected by the user. The user interface (UI) displays the dialogue scenario visually and audibly, allowing the user to begin the dialogue. The input here is the dialogue scenario sent from the server, and the output is the dialogue scenario displayed on the user interface.
[1496] Step 4:
[1497] The user inputs a question or response into the terminal. Speech recognition technology is used to convert the user's voice input into text, and the text is sent to the server. The input is the user's voice, and the output is the user's question or response converted into text.
[1498] Step 5:
[1499] The server recognizes the questions and responses entered by the user and generates an appropriate response. It uses a generative AI model to perform natural language processing and generate an appropriate response to the user's question. The input is the text received in step 4, and the output is the generated response.
[1500] Step 6:
[1501] The terminal outputs the response received from the server to the user as voice. It uses speech synthesis technology (pyttsx3) to convey information in a way that is easy for the user to understand. The input is the text of the response sent from the server, and the output is the voice response.
[1502] Step 7:
[1503] The device recognizes the user's emotions and sends the information to the server. Emotion recognition uses OpenCV and a pre-trained emotion recognition model. The input is the user's facial image and voice, and the output is the recognized emotion information.
[1504] Step 8:
[1505] The server adjusts the content and tone of the response based on the emotional information. It changes the tone and content of the response depending on the emotion expressed by the user, generating a more personalized response. The input is the emotional information received in step 7, and the output is the adjusted response.
[1506] Step 9:
[1507] The device outputs the adjusted response to the user as voice, again using speech synthesis technology to provide an adapted response to the user. The input is the adjusted response text sent from the server, and the output is the voice response.
[1508] Step 10:
[1509] The terminal visually displays the beer brewing process in a virtual environment. Using virtual reality and augmented reality technology, the user can visually experience the beer brewing process. The input is visual data sent from the server, and the output is the beer brewing process displayed in the virtual environment.
[1510] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1511] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1512] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1513] [Fourth embodiment]
[1514] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1515] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1516] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1517] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1518] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1519] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1520] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1521] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1522] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1523] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1524] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1525] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1526] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1527] This invention relates to a system that allows users to deepen their language learning and intercultural understanding through the beer brewing process. The system involves three components: a server, a terminal, and a user, and clarifies the roles of each component.
[1528] Server Processing
[1529] The server serves as the core of this system and provides the following specific functions:
[1530] 1. Receiving user information:
[1531] The server receives language selection information and basic information input by the user through the terminal.
[1532] 2. Generating dialogue scenarios:
[1533] The server generates a dialogue scenario based on the beer brewing process. This dialogue scenario is structured appropriately based on the language selected by the user. Advanced generative AI such as ChatGPT is used to achieve natural dialogue.
[1534] 3. Response Generation:
[1535] The server receives questions and responses from users and uses AI technology to generate appropriate responses.
[1536] 4. Database Update:
[1537] The server periodically adds the latest information about the beer production process and new dialogue scenarios to the database and provides them to users.
[1538] Terminal handling
[1539] The terminal provides the following specific functions as an interface with the user:
[1540] 1. User Interface:
[1541] The terminal acts as a virtual tour guide, displaying each stage of the beer production process in the user's language of choice.
[1542] 2. Speech recognition and output:
[1543] The terminal converts the user's voice input into text and also synthesizes the text received from the server into speech for the user to hear.
[1544] 3. Data synchronization:
[1545] The device periodically synchronizes learning progress with the server, providing the user with the most up-to-date information.
[1546] User Behavior
[1547] The user performs the following specific actions through this system:
[1548] 1. Initial Settings and Selections:
[1549] Launch the app, select the language you want to learn, and enter your basic information.
[1550] 2. Selection of beer production process:
[1551] Within the app, users select the stage of the beer-making process they want to learn about (e.g., fermentation).
[1552] 3. Execute the dialogue:
[1553] The user interacts with the system at each stage of the process to learn. For example, in response to a question about the fermentation process, "Tell me how alcohol is produced during fermentation," the server responds, "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide."
[1554] Specific examples
[1555] Example: A user is learning about the beer fermentation process.
[1556] 1. Initial Setup:
[1557] The user launches the app, selects Japanese, and enters basic information. The device sends the information to the server, which stores it in a database and generates a learning plan optimized for the user.
[1558] 2. Process Selection:
[1559] The user selects "fermentation process" and the terminal sends this information to the server, which then generates a corresponding dialogue scenario and sends it to the terminal.
[1560] 3. Start a conversation:
[1561] The device starts the dialogue scenario by asking, "Do you know how alcohol is produced during the fermentation process?" The user responds, "I don't know. Please tell me." The device sends this to the server, which responds, "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide," which the device then relays in voice.
[1562] 4. Continuing the dialogue:
[1563] The user asks a further question (e.g., "What types of yeast are there?"). The device recognizes the question and sends it to the server, which generates an appropriate response, which the device then speaks to the user.
[1564] This system allows users to deepen their language learning and intercultural understanding through natural dialogue while using the beer brewing process as a subject, thereby increasing their motivation to learn.
[1565] The processing flow will be explained below.
[1566] Step 1:
[1567] The user launches the app. The device displays a startup screen, prompting the user to select a language and enter basic information.
[1568] Step 2:
[1569] The user inputs the desired language and basic information, which the terminal then sends to the server.
[1570] Step 3:
[1571] The server stores the received user information in a database, generates an optimal learning plan for the user, and sends it to the device.
[1572] Step 4:
[1573] Based on the learning plan received from the server, the device displays a selection screen for the beer brewing process, prompting the user to select the process they wish to learn.
[1574] Step 5:
[1575] The user selects "Fermentation process." The terminal transmits the user's selection to the server.
[1576] Step 6:
[1577] The server generates a dialogue scenario corresponding to the "fermentation process." It uses generative AI such as ChatGPT to build the dialogue scenario in the language selected by the user.
[1578] Step 7:
[1579] The server transmits the generated dialogue scenario to the terminal.
[1580] Step 8:
[1581] The device displays a dialogue scenario to the user and also starts voice output, asking questions such as, "Do you know how alcohol is produced during the fermentation process?"
[1582] Step 9:
[1583] The user responds to the question (e.g., "I don't understand. Please tell me."). The device recognizes the user's voice, converts it into text, and sends it to the server.
[1584] Step 10:
[1585] The server generates an appropriate response based on the user's response, using AI techniques to create answers in natural language (e.g., "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide.").
[1586] Step 11:
[1587] The server sends the generated response to the terminal.
[1588] Step 12:
[1589] The terminal synthesizes the response from the server into voice and plays it back to the user.
[1590] Step 13:
[1591] The user continues to ask questions (e.g., "What kinds of yeast are there?"). The device converts the question into text using voice recognition and sends it to the server.
[1592] Step 14:
[1593] The server generates a response to the new question (e.g., "There are many different types of yeast, including ale yeast and lager yeast.") and sends it to the terminal.
[1594] Step 15:
[1595] The device synthesizes the response into voice and plays it back to the user. This process is repeated to advance learning in an interactive format.
[1596] Step 16:
[1597] The device periodically sends the user's learning progress to the server, which records the progress and uses it to optimize the learning content for the next session.
[1598] Step 17:
[1599] The server periodically updates the database, adding new dialogue scenarios and the latest information on beer production, and the terminal provides the new information to the user.
[1600] Example 1
[1601] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1602] Conventional language learning systems struggle to sustain users' interest and attention, leading to a decline in motivation to learn. Furthermore, most language learning systems simply focus on language acquisition, failing to incorporate cultural background and practical knowledge, making it difficult to deepen intercultural understanding. Furthermore, responses to user questions are mechanical, making it difficult to achieve natural dialogue. This tends to reduce learning efficiency and user satisfaction.
[1603] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1604] In this invention, the server includes means for receiving language selection information and basic information input by the user, means for generating a dialogue scenario related to a manufacturing process such as a fermentation process, means for providing the generated dialogue scenario in the user's selected language, means for recognizing questions and responses input by the user and generating appropriate replies, and means for outputting the generated replies to the user. This makes it possible to deepen language learning and intercultural understanding in a natural dialogue format while maintaining the user's interest, thereby improving learning efficiency and user satisfaction.
[1605] "User" refers to an individual who uses this system to input information and interact in order to deepen language learning and intercultural understanding.
[1606] "Language selection information" is information that a user inputs to the system when selecting the language they wish to learn.
[1607] "Basic information" refers to personal information such as name, age, and interests that a user provides to the system.
[1608] The "server" is the core of this system, and refers to a computer system that processes information received from the user and generates dialogue scenarios and responses.
[1609] The "manufacturing process" specifically refers to each stage of production activities, such as the fermentation process of beer, and dialogue scenarios are generated based on this theme.
[1610] A "dialogue scenario" is a planned conversation content for progressing learning and dialogue, which is generated based on the manufacturing process selected by the user.
[1611] "Response" refers to the words or documents generated by the system in response to a user's question or response.
[1612] "Speech recognition" is a technology that converts a user's voice input into text data.
[1613] "Speech synthesis" is a technology that converts text data into speech and reads it out to the user.
[1614] "Data Server" refers to the computer system used to record and store users' learning progress and latest information.
[1615] "Database" refers to a data storage system for systematically storing collected information and generated scenarios.
[1616] This invention is a system that allows users to deepen their language learning and intercultural understanding through the manufacturing process. This system involves three components: a server, a terminal, and a user, and each component clearly divides its role.
[1617] Server Processing
[1618] The server serves as the core of this system and provides the following specific functions:
[1619] 1. Receiving user information:
[1620] The server receives the language selection information and basic information entered by the user through the terminal, and records the information in a user profile database using specific software such as an HTTP server (e.g., Nginx) and a database server (e.g., MySQL).
[1621] 2. Generating dialogue scenarios:
[1622] The server generates a dialogue scenario based on a manufacturing process (e.g., a fermentation process). Here, generative AI (e.g., ChatGPT) is used to create a natural and interactive dialogue scenario. For example, the following prompts are used:
[1623] The user wants to learn about the "fermentation process." Generate a clear explanation of the fermentation process.
[1624] 3. Response Generation:
[1625] The server receives questions and responses from the user and uses generative AI to generate appropriate responses. For example, if the user asks, "Tell me how alcohol is produced during fermentation," the server sends the following prompt to the generative AI:
[1626] User: How is alcohol produced during fermentation?
[1627] The generative AI replies, "During the fermentation process, yeast converts sugars into alcohol and carbon dioxide."
[1628] 4. Database Update:
[1629] The server periodically adds the latest information about the manufacturing process and new interaction scenarios to the database and provides them to users. The latest information is collected through web crawlers and external APIs and added to the database management system (DBMS).
[1630] Terminal handling
[1631] The terminal provides the following specific functions as an interface with the user:
[1632] 1. User Interface:
[1633] The device acts as a virtual tour guide, displaying each stage of the manufacturing process in the user's language of choice, and uses cross-platform frameworks such as React Native for its software.
[1634] 2. Speech recognition and output:
[1635] The device uses speech recognition technology (e.g., Google Cloud Speech-to-Text API) to convert the user's voice input into text, and also converts the text received from the server into speech using speech synthesis technology (e.g., Amazon Polly) and provides it to the user.
[1636] 3. Data synchronization:
[1637] The device periodically synchronizes its learning progress with the server to provide the user with the latest information, using a long-lived connection over WebSocket or HTTP.
[1638] User Behavior
[1639] The user performs the following specific actions through this system:
[1640] 1. Initial Settings and Selections:
[1641] Users launch the app, select the language they want to learn, and enter basic information. The device sends this information to the server, which stores it in a database. The server then generates a personalized learning plan for the user.
[1642] 2. Selection of beer production process:
[1643] Within the app, the user selects the stage of the manufacturing process they want to learn about (e.g., the fermentation stage). The device sends this information to the server, which then generates a corresponding dialogue scenario and sends it to the device.
[1644] 3. Execute the dialogue:
[1645] The user interacts with the system at each stage of the process to learn: for example, if a user asks, "Tell me how alcohol is produced during fermentation," the server responds, "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide."
[1646] 4. Continuing the dialogue:
[1647] If the user asks a further question (e.g., "What types of yeast are there?"), the device recognizes this question and sends it to the server, which generates an appropriate response, which the device then speaks to the user.
[1648] This system allows users to deepen their language learning and intercultural understanding through a natural dialogue format while using manufacturing processes as a subject, thereby increasing their motivation to learn.
[1649] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1650] Step 1: User launches the app
[1651] Input: A user taps on a device to launch an app.
[1652] Output: A language selection screen will appear.
[1653] What happens: When a user launches the app, the device loads the app's initial screen. The app is built with a cross-platform framework such as React Native and prompts the user to select a language.
[1654] Step 2: User enters language and basic information
[1655] Input: Enter your language preference and basic information (name, age, interests, etc.).
[1656] Output: The entered information is sent to the server and stored in a database.
[1657] Specific operation: The user enters the language and basic information in the input form and presses the "Submit" button. The terminal sends this as an HTTP POST request to the server, which receives the request via Nginx and stores it in the MySQL database.
[1658] Step 3: The server generates the lesson plan
[1659] Input: User's language preference and basic information.
[1660] Output: User-optimized learning plan data.
[1661] Specific operation: The server generates an optimal learning plan based on the received user information, referencing the data in the database. The generated learning plan is returned to the device in JSON format.
[1662] Step 4: User selects stage in manufacturing process
[1663] Input: The stage of the manufacturing process you want to learn about (e.g., fermentation process).
[1664] Output: Stage selection information is sent to the server.
[1665] Specific operation: When a user selects "Fermentation process" in the app, the device sends this information in JSON format to the server as an HTTP POST request. The server receives this information.
[1666] Step 5: The server generates the dialogue scenario
[1667] Input: Manufacturing process stage information.
[1668] Output: Dialogue scenario data.
[1669] Specific operation: The server sends prompts to generative AI models such as ChatGPT based on the manufacturing process. For example, it sends the following prompts:
[1670] The user wants to learn about the "fermentation process." Generate a clear explanation of the fermentation process.
[1671] The server receives the scenario generated by ChatGPT and returns it to the terminal in JSON format.
[1672] Step 6: The device presents the dialogue scenario
[1673] Input: Dialogue scenario data.
[1674] Output: Presents the dialogue scenario to the user in voice and text.
[1675] Specific operation: The device analyzes the received dialogue scenario and presents it to the user in both voice and text format. It converts the text into speech using speech synthesis technology (e.g., Amazon Polly).
[1676] Step 7: User enters question and response
[1677] Input: User voice or text input.
[1678] Output: The entered questions and responses are sent to the server in text format.
[1679] Specific operation: The user responds with something like "I don't understand. Please tell me." The device uses the Google Cloud Speech-to-Text API to convert the speech into text and send it to the server.
[1680] Step 8: The server generates a response
[1681] Input: The user's question and response.
[1682] Output: The generated response data.
[1683] How it works: The server receives the user's question and sends the following prompt to the generative AI (e.g. ChatGPT):
[1684] User: How is alcohol produced during fermentation?
[1685] The server receives the response generated by ChatGPT: "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide." and sends it to the terminal.
[1686] Step 9: The device presents a response
[1687] Input: The generated response data.
[1688] Output: Presents the response to the user in audio and text.
[1689] Specific operation: The device analyzes the response received from the server and uses speech synthesis technology to present it to the user in the form of voice and text.
[1690] This series of steps allows users to deepen their language learning and intercultural understanding through a natural, interactive way during the manufacturing process.
[1691] (Application example 1)
[1692] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1693] Conventional language learning systems have limited content that attracts users' interest, making it difficult to maintain motivation. In particular, they are limited to learning grammar and vocabulary, with few opportunities to learn cultural background or practical applications. Another problem is that the lack of interactive learning makes it difficult to improve practical language skills.
[1694] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1695] In this invention, the server includes means for receiving language selection information and basic information input by a user, means for generating a dialogue scenario related to the beer brewing process, means for providing the generated dialogue scenario in a language selected by the user, means for recognizing questions and responses input by the user and generating appropriate responses, means for outputting the generated responses to the user, means for visually representing the beer brewing process in a virtual environment, and means for the user to learn a language through dialogue in the virtual environment. This enables the user to deepen practical language learning and intercultural understanding through interactive dialogue in the virtual environment using the beer brewing process as a subject.
[1696] "Language selection information and basic information entered by the user" refers to the language used for learning that the user selects for the system and basic personal information.
[1697] The "means for generating a dialogue scenario" refers to a system component that has the function of automatically creating a dialogue scenario with the user based on information about the beer brewing process.
[1698] "Means for providing in a language selected by the user" refers to a system component that has the function of translating the generated dialogue scenario into a specific language selected by the user and presenting information in that language.
[1699] "Means for recognizing questions and responses and generating appropriate responses" refers to a system component that has the function of recognizing and analyzing questions and responses entered by users in voice or text format and automatically generating appropriate responses to them.
[1700] "Means for outputting the generated response to the user" refers to a system component that has the function of presenting the response generated by the system to the user in the form of voice or text.
[1701] "Means for visually representing the beer production process in a virtual environment" refers to a component of a system that uses virtual reality technology to visually recreate the beer production process and provide users with an interactive experience.
[1702] "Means for users to learn a language through interaction in a virtual environment" refers to a component of a system that has the function of allowing users to learn a language through interaction in an environment created using virtual reality technology.
[1703] This invention relates to a system that aims to deepen language learning and intercultural understanding through the beer brewing process. Specifically, it utilizes virtual reality technology and generative AI models to provide users with an interactive learning experience.
[1704] Server Processing
[1705] The server serves as the core of this system and provides the following specific functions:
[1706] 1. Receiving user information:
[1707] The server receives the language selection information and basic information input by the user through the terminal and stores them in a database.
[1708] 2. Dialogue scenario generation:
[1709] The server generates a dialogue scenario based on the beer brewing process, and the dialogue scenario is structured appropriately based on the language selected by the user. Advanced generative AI models such as ChatGPT are used to achieve natural dialogue.
[1710] 3. Response Generation:
[1711] The server receives questions and responses from users and uses generative AI technology to generate appropriate responses.
[1712] 4. Data synchronization:
[1713] The server records the user's learning progress and periodically adds the latest information and new dialogue scenarios to the database and provides them to the user.
[1714] Terminal handling
[1715] The terminal provides the following specific functions as an interface with the user:
[1716] 1. User Interface:
[1717] It acts as a virtual tour guide, showing each stage of the beer production process in the user's language of choice.
[1718] 2. Speech Recognition and Synthesis:
[1719] The system converts the user's voice input into text, synthesizes the text received from the server into speech, and plays it back to the user.
[1720] 3. Virtual Environment:
[1721] Users use smart glasses or a head-mounted display (HMD) to visually experience the beer production process in a virtual environment.
[1722] User Behavior
[1723] The user performs the following specific actions through this system:
[1724] 1. Initial Settings and Selections:
[1725] Launch the app, select the language you want to learn, and enter some basic information. The device then sends the information to the server, which stores it in a database.
[1726] 2. Selection of beer production process:
[1727] The user selects the stage of the beer production process (e.g., fermentation) they want to learn about in the virtual environment. The device sends this information to the server, which then generates a corresponding dialogue scenario.
[1728] 3. Execute the dialogue:
[1729] For example, in response to a question about the fermentation process, "Tell me how alcohol is produced during fermentation," the server responds, "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide."
[1730] Specific examples
[1731] Consider a case where a user is learning about the beer fermentation process.
[1732] 1. Initial Setup:
[1733] The user launches the app, selects Japanese, and enters basic information. The device sends the information to the server, which stores it in a database and generates a learning plan optimized for the user.
[1734] 2. Process Selection:
[1735] The user selects "fermentation process" and the terminal sends this information to the server, which then generates a corresponding dialogue scenario and sends it to the terminal.
[1736] 3. Start a conversation:
[1737] The device starts the dialogue scenario by asking, "Do you know how alcohol is produced during fermentation?" The user responds, "I don't know. Please tell me." The device sends this to the server, which replies, "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide," which the device then relays.
[1738] 4. Further dialogue:
[1739] The user then asks a further question (e.g., "What kinds of yeast are there?"). The device recognizes the question and sends it to the server, which generates an appropriate response, which the device then relays to the user via voice.
[1740] Prompt Sentence Examples
[1741] Below is an example of a prompt sentence to input to the generative AI model.
[1742] User Question: How is alcohol produced during fermentation?
[1743] Correct Response: During the fermentation process, yeast converts sugars into alcohol and carbon dioxide.
[1744] This system allows users to deepen their practical language learning and intercultural understanding through interactive dialogue in a virtual environment using the beer brewing process as a subject.
[1745] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1746] Step 1:
[1747] Entering and receiving user information
[1748] The user launches the application and inputs the language they want to learn and basic information through smart glasses or a head-mounted display (HMD). The device receives this information and sends it to the server. The input data is the user's language of choice and personal information, and the output is the user's basic information and learning plan, which are stored in a database. The server receives this data and stores it in a database.
[1749] Step 2:
[1750] Dialogue scenario generation
[1751] The server generates a dialogue scenario based on the user's selected language and information about the beer brewing process. It uses a generative AI model such as ChatGPT to create a scenario for natural dialogue. The input data are details of the beer brewing process and the user's selected language, and the output is the generated dialogue scenario. The server generates the dialogue scenario and sends it to the device.
[1752] Step 3:
[1753] Building a virtual environment
[1754] Based on the received dialogue scenario, the device creates a virtual environment in the smart glasses or head-mounted display (HMD) worn by the user. The input data is the dialogue scenario and virtual environment setting information sent from the server, and the output is a visual virtual environment provided to the user. The device displays each stage of the beer brewing process in a virtual space.
[1755] Step 4:
[1756] Starting a conversation
[1757] The user selects a specific stage of the beer brewing process in the virtual environment and begins a dialogue. The device sends this selection information to the server, which then prepares a corresponding dialogue scenario. The input data is the information about the stage selected by the user, and the output is the dialogue scenario prepared by the server. The device plays the scenario aloud and presents questions to the user.
[1758] Step 5:
[1759] Speech Recognition and Processing
[1760] The user responds to the question by voice. The device converts the user's voice into text data using voice recognition software and sends the text data to the server. The input data is the user's voice data, and the output is text data. The server receives this data and uses it in the next step.
[1761] Step 6:
[1762] Response Generation
[1763] The server uses a generative AI model to generate an appropriate response based on the received text data. The input data is the text data received from the user, and the output is the generated response text. The server generates this response and sends it to the device.
[1764] Step 7:
[1765] Response output
[1766] The terminal converts the response text received from the server into speech using speech synthesis software and presents the response to the user. The input data is the text data received from the server, and the output is a voice message provided to the user. As a specific example, if the user asks, "Please tell me how alcohol is produced during the fermentation process," the terminal will respond by voice, "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide."
[1767] Step 8:
[1768] Record your learning progress
[1769] The device records the user's learning progress and sends it to the server, which stores this information in a database and updates the user's optimized learning plan. The input data is the user's learning progress information, and the output is the updated learning information stored in the database.
[1770] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1771] This invention is a system that aims to help users deepen their language learning and intercultural understanding through the beer brewing process, and by incorporating an emotion engine that recognizes the user's emotions, it provides a more personalized learning experience. The system involves three parties: a server, a terminal, and a user, and clarifies the roles of each part.
[1772] Server Processing
[1773] The server serves as the core of this system and provides the following specific functions:
[1774] 1. Receiving user information:
[1775] The server receives language selection information and basic information input by the user through the terminal.
[1776] 2. Generating dialogue scenarios:
[1777] The server generates a dialogue scenario based on the beer brewing process. This dialogue scenario is structured appropriately based on the language selected by the user. Advanced generative AI such as ChatGPT is used to achieve natural dialogue.
[1778] 3. Response Generation:
[1779] The server receives questions and responses from users and uses AI technology to generate appropriate responses.
[1780] 4. Emotional Information Processing:
[1781] The server receives the user's emotional information sent from the terminal and adjusts the content and tone of the response to the user based on this information.
[1782] 5. Database Update:
[1783] The server periodically adds the latest information about the beer production process and new dialogue scenarios to the database and provides them to users.
[1784] Terminal handling
[1785] The terminal provides the following specific functions as an interface with the user:
[1786] 1. User Interface:
[1787] The terminal acts as a virtual tour guide, displaying each stage of the beer production process in the user's language of choice.
[1788] 2. Speech recognition and output:
[1789] The terminal converts the user's voice input into text and also synthesizes the text received from the server into speech for the user to hear.
[1790] 3. Emotion recognition:
[1791] The device uses voice input and facial recognition technology to recognize the user's emotions and sends that information to a server.
[1792] 4. Data synchronization:
[1793] The device periodically synchronizes learning progress with the server, providing the user with the most up-to-date information.
[1794] User Behavior
[1795] The user performs the following specific actions through this system:
[1796] 1. Initial Settings and Selections:
[1797] Launch the app, select the language you want to learn, and enter your basic information.
[1798] 2. Selection of beer production process:
[1799] Within the app, users select the stage of the beer-making process they want to learn about (e.g., fermentation).
[1800] 3. Execute the dialogue:
[1801] The user interacts with the system at each stage of the process to learn. For example, in response to a question about the fermentation process, "Tell me how alcohol is produced during fermentation," the server responds, "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide."
[1802] Specific examples
[1803] Example: A user is learning about the beer fermentation process.
[1804] 1. Initial Setup:
[1805] The user launches the app, selects Japanese, and enters basic information. The device sends the information to the server, which stores it in a database and generates a learning plan optimized for the user.
[1806] 2. Process Selection:
[1807] The user selects "fermentation process" and the terminal sends this information to the server, which then generates a corresponding dialogue scenario and sends it to the terminal.
[1808] 3. Start a conversation:
[1809] The device starts the dialogue scenario by asking, "Do you know how alcohol is produced during the fermentation process?" The user responds, "I don't know. Please tell me." The device sends this to the server, which responds, "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide," which the device then relays in voice.
[1810] 4. Emotion Recognition and Processing:
[1811] If the user shows an emotional reaction to a question (e.g., a questioning or surprised expression), the device recognizes that emotion and sends it to the server. The server then adjusts the tone and content of the response based on the emotional information, and generates an appropriate response and sends it to the device.
[1812] 5. Continuing the dialogue:
[1813] The user asks a further question (e.g., "What types of yeast are there?"). The device recognizes the question and sends it to the server, which generates an appropriate response, which the device then speaks to the user.
[1814] In this way, users can deepen their language learning and intercultural understanding through natural dialogue while learning about the beer brewing process. The introduction of an emotion engine provides responses that correspond to the user's emotions, creating a more personalized learning experience.
[1815] The processing flow will be explained below.
[1816] Step 1:
[1817] The user launches the app. The device displays a startup screen, prompting the user to select a language and enter basic information.
[1818] Step 2:
[1819] The user inputs the desired language and basic information, which the terminal then sends to the server.
[1820] Step 3:
[1821] The server stores the received user information in a database, generates an optimal learning plan for the user, and sends it to the device.
[1822] Step 4:
[1823] Based on the learning plan received from the server, the device displays a selection screen for the beer brewing process, prompting the user to select the process they wish to learn.
[1824] Step 5:
[1825] The user selects "Fermentation process." The terminal transmits the user's selection to the server.
[1826] Step 6:
[1827] The server generates a dialogue scenario corresponding to the "fermentation process." It uses generative AI such as ChatGPT to build the dialogue scenario in the language selected by the user.
[1828] Step 7:
[1829] The server transmits the generated dialogue scenario to the terminal.
[1830] Step 8:
[1831] The device displays the dialogue scenario to the user and also starts voice output, asking the question, "Do you know how alcohol is produced during the fermentation process?"
[1832] Step 9:
[1833] The user responds to the question (e.g., "I don't understand. Please tell me."). The device recognizes the user's voice, converts it into text, and sends it to the server.
[1834] Step 10:
[1835] The server generates an appropriate response based on the user's response, using AI techniques to create answers in natural language (e.g., "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide.").
[1836] Step 11:
[1837] The server sends the generated response to the terminal.
[1838] Step 12:
[1839] The terminal synthesizes the response from the server into voice and plays it back to the user.
[1840] Step 13:
[1841] The user asks a follow-up question (e.g., "What kinds of yeast are there?"). The device uses voice recognition to convert the question into text and sends it to the server.
[1842] Step 14:
[1843] The server generates a response to the new question (e.g., "There are many different types of yeast, including ale yeast and lager yeast.") and sends it to the terminal.
[1844] Step 15:
[1845] The device synthesizes the response into speech and plays it back to the user.
[1846] Step 16:
[1847] The emotion engine recognizes the tone of voice and facial expressions used during the user's questions and responses, and the device analyzes this and transmits the user's emotional information to the server.
[1848] Step 17:
[1849] Based on the received emotional information, the server adjusts the next response to match the tone and content of the user's feelings.
[1850] Step 18:
[1851] The device will provide users with responses tailored to their emotions, enabling more personalized interactions.
[1852] Step 19:
[1853] The device periodically sends the user's learning progress to the server, which records the progress and uses it to optimize the learning content for the next session.
[1854] Step 20:
[1855] The server periodically updates the database, adding new dialogue scenarios and the latest information on beer production, and the terminal provides the new information to the user.
[1856] Example 2
[1857] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1858] Traditional language learning systems struggled to provide a personalized experience based on users' interests and emotions. They also lacked dynamic learning content that incorporated the latest information. As a result, it was difficult to maintain learners' motivation, and effective learning could not be expected.
[1859] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1860] In this invention, the server includes means for receiving language selection information and basic information input by a user, means for generating a dialogue scenario related to the beer brewing process, means for providing the generated dialogue scenario in the user's language of choice, means for recognizing questions and responses input by the user and generating appropriate replies, means for analyzing the user's emotional state and adjusting the content and tone of the replies, means for outputting the generated replies to the user, means for adding the latest information to the database, and means for periodically updating the information in the database and providing it to the user. This provides a personalized learning experience based on the user's interests and emotions, and enables dynamic learning incorporating the latest information, thereby enabling effective language learning.
[1861] The "means for receiving language selection information and basic information input by the user" refers to an interface for receiving the language selected by the user through the application and basic information such as name, language level, and learning purpose.
[1862] The "means for generating dialogue scenarios related to the beer production process" is a system that uses a generative AI model or the like to generate dialogue-style scenarios based on each stage of beer production.
[1863] The "means for providing the generated dialogue scenario in a language selected by the user" refers to a means for displaying or audibly outputting the generated dialogue scenario in a language selected by the user.
[1864] "Means for recognizing questions and responses entered by users and generating appropriate responses" refers to a system that utilizes AI technology to recognize questions and responses from users in text or voice and generate responses to them.
[1865] "Means for analyzing the user's emotional state and adjusting the content and tone of the response" refers to means for detecting the user's emotional state using voice analysis or facial recognition technology and adjusting the content and tone of the response according to that emotion.
[1866] The "means for outputting the generated response to the user" refers to a means for providing the generated response from the server to the user in the form of voice or text.
[1867] The "means for adding the latest information to the database" is a means for periodically collecting the latest information on the beer production process and adding it to the database.
[1868] "Means for periodically updating the information in the database and providing it to users" refers to means for keeping the information in the database up to date and providing that information to users.
[1869] This invention is a system that aims to help users deepen their language learning and intercultural understanding through the beer brewing process, and by incorporating an emotion engine that recognizes the user's emotions, it provides a more personalized learning experience. The system involves three parties: a server, a terminal, and a user, and clarifies the roles of each part.
[1870] Specific server processing
[1871] The server serves as the core of this system and provides the following specific functions:
[1872] 1. Receiving user information:
[1873] The server receives language selection information and basic information entered by the user through the terminal, which is stored in a database and used to create a user profile.
[1874] 2. Generating dialogue scenarios:
[1875] The server generates a dialogue scenario based on the beer brewing process. This dialogue scenario is structured appropriately based on the language selected by the user. Natural dialogue is achieved using generative AI models such as ChatGPT.
[1876] 3. Response Generation:
[1877] The server receives questions and responses from users and uses AI technology to generate appropriate responses, allowing users to learn specialized knowledge about beer brewing.
[1878] 4. Emotional Information Processing:
[1879] The server receives the user's emotional information sent from the terminal and adjusts the content and tone of the response to the user based on this information.
[1880] 5. Database Update:
[1881] The server periodically adds the latest information about the beer production process and new dialogue scenarios to the database and provides them to users.
[1882] Specific processing of the terminal
[1883] The terminal provides the following specific functions as an interface with the user:
[1884] 1. User Interface:
[1885] The terminal acts as a virtual tour guide, displaying each stage of the beer production process in the user's language of choice.
[1886] 2. Speech recognition and output:
[1887] The terminal converts the user's voice input into text and also synthesizes the text received from the server into speech for the user to hear.
[1888] 3. Emotion recognition:
[1889] The device uses voice input and facial recognition technology to recognize the user's emotions and sends that information to a server.
[1890] 4. Data synchronization:
[1891] The device periodically synchronizes learning progress with the server, providing the user with the most up-to-date information.
[1892] Specific user actions
[1893] The user will perform the following specific actions through this system:
[1894] 1. Initial Settings and Selections:
[1895] Users launch the app, select the language they want to learn, and enter basic information. The device sends this information to the server, which stores it in a database and generates a personalized learning plan.
[1896] 2. Selection of beer production process:
[1897] Within the app, users select the stage of the beer brewing process they want to learn about (e.g., the fermentation stage). The device sends this information to the server, which then generates a corresponding dialogue scenario and sends it to the device.
[1898] 3. Execute the dialogue:
[1899] The user interacts with the system at each stage of the process. For example, in response to a question about the fermentation process, "Tell me how alcohol is produced during fermentation," the server responds with, "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide." For example, the user can ask a question using the following prompt:
[1900] "Please explain the fermentation process of beer in detail."
[1901] What types of yeast are there?
[1902] summary
[1903] The system is designed to enable users to deepen their language learning and intercultural understanding through natural dialogue while learning about the beer brewing process. The system incorporates an emotion engine to provide responses based on the user's emotions, creating a more personalized learning experience. The system also adds the latest information to the database and provides it to users, ensuring they are constantly learning new knowledge.
[1904] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1905] Program processing flow
[1906] Step 1: Enter and submit user information
[1907] Specific behavior:
[1908] 1. Input: The user launches the app, selects Japanese or another language they want to learn, and enters basic information such as their name, language level, and learning purpose.
[1909] 2. Data processing: The device collects this information and converts it into the appropriate format.
[1910] 3. Send: The device creates and sends an API request to send the formatted information to the server.
[1911] 4. Output: The server saves the user's basic information and language preference in a database and creates a user profile.
[1912] Step 2: Generate dialogue scenarios
[1913] Specific behavior:
[1914] 1. Input: The server retrieves information about the beer production process from a database.
[1915] 2. Data processing: Based on the information obtained, the server converts the prompt text into a format suitable for sending to the generative AI model.
[1916] 3. Generation: A generative AI model like ChatGPT generates a dialogue scenario based on the prompt.
[1917] 4. Output: The server receives the generated dialogue scenario, converts it into an appropriate format, and sends it to the terminal to be presented in the language selected by the user.
[1918] Step 3: Handling User Questions and Responses
[1919] Specific behavior:
[1920] 1. Input: The user enters questions and responses by voice or text at each stage of the process, for example, "Tell me how alcohol is produced during fermentation."
[1921] 2. Data processing: The device uses voice recognition technology to convert the voice into text and sends the text data to the server.
[1922] 3. Generation: The server analyzes the user's question, sends it as a prompt to the generative AI model, and generates an appropriate response.
[1923] 4. Output: The server receives the generated response and sends it to the terminal, which then provides the response to the user using voice synthesis technology.
[1924] Step 4: Recognizing and processing emotional information
[1925] Specific behavior:
[1926] 1. Input: The user asks questions and responds to the device, and emotions are input from facial expressions and tone of voice.
[1927] 2. Data processing: The device uses voice analysis and facial recognition technology to analyze the user's emotional state, for example, detecting expressions of doubt or surprise.
[1928] 3. Send: The device sends the emotion analysis results to the server.
[1929] 4. Processing: The server adjusts the tone and content of the response based on the emotional information, generates an appropriate response, and sends it to the device.
[1930] 5. Output: The device communicates the adjusted response received from the server to the user via voice synthesis.
[1931] Step 5: Update the database
[1932] Specific behavior:
[1933] 1. Input: The server retrieves up-to-date information about the beer production process from an external resource.
[1934] 2. Data processing: The server converts the acquired information into a format for adding to the database.
[1935] 3. Add: The server adds new information to the database and updates existing information.
[1936] 4. Output: Based on the updated information, the server generates a new dialogue scenario and sends it to the terminal to provide the user with up-to-date information.
[1937] (Application example 2)
[1938] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1939] The purpose of this invention is to provide a language learning system based on the beer brewing process that can personalize the user's learning experience and provide an interactive and effective learning environment. In particular, the invention aims to deepen the user's interest and understanding in learning by recognizing the user's emotions and adjusting the content and tone of responses accordingly, and to provide an experience based on actual procedures and culture by visually displaying the beer brewing process in a virtual environment.
[1940] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving language selection information and basic information input by the user, means for generating a dialogue scenario related to the beer brewing process, means for providing the generated dialogue scenario in the user's language, means for recognizing questions and responses input by the user and generating appropriate replies, means for outputting the generated replies to the user, means for recognizing emotions and adjusting the content and tone of the responses, and means for visually displaying the beer brewing process in a virtual environment. This enables users to deepen their language learning and intercultural understanding through a natural dialogue format through the beer brewing process. Furthermore, the use of emotion recognition and a virtual environment makes the learning experience more interactive and personalized.
[1941] "Means for receiving user-entered language selection information and basic information" refers to a device for entering the user's preferred language and typical profile information (e.g., name, age, affiliation, etc.) and the function for transmitting that information to the system.
[1942] "Means for generating dialogue scenarios related to the beer production process" refers to a function for generating scenarios for dialogue in a language selected by the learner based on each stage of beer production.
[1943] "Means for providing the generated dialogue scenario in a language selected by the user" refers to an interface and device for translating the generated dialogue scenario into a language selected by the user and effectively providing it.
[1944] "Means of recognizing questions and responses entered by the user and generating appropriate responses" refers to a function that uses voice recognition technology and natural language processing technology to understand the content entered by the user and automatically generate appropriate responses based on that.
[1945] "Means for outputting the generated response to the user" refers to an output device or software that conveys the system-generated response to the user in audio or text form.
[1946] "Means for recognizing emotions and adjusting the content and tone of responses" refers to a function that analyzes emotions from the user's facial expressions and voice, and optimizes the system's response content and speaking style according to those emotions.
[1947] "Means for visually displaying the beer production process in a virtual environment" refers to a function that uses virtual reality or augmented reality technology to display the production process so that the user can visually understand each stage of beer production.
[1948] This invention is a system that aims to help users deepen their language learning and intercultural understanding through the beer brewing process, and by incorporating an emotion engine that recognizes the user's emotions, it provides a more personalized learning experience. Here, we will explain in detail how to specifically implement the invention.
[1949] System Overview
[1950] This system is realized by the collaboration of three parties: the server, the terminal, and the user, each fulfilling their respective roles. The server is responsible for all main processing, while the terminal functions as an interface with the user, allowing the user to learn through dialogue.
[1951] Program processing
[1952] Server Processing
[1953] 1. Receiving user information:
[1954] The server receives the language selection information and basic information entered by the user through the terminal, and this function is used to securely store the entered data and generate the necessary study plan.
[1955] 2. Generating dialogue scenarios:
[1956] The server generates a dialogue scenario based on the beer brewing process. This dialogue scenario is structured appropriately based on the language selected by the user. It uses OpenAI's generative AI model to achieve natural dialogue.
[1957] 3. Response Generation:
[1958] The server receives questions and responses from users and uses natural language processing technology to generate appropriate responses, using the latest generative AI models.
[1959] 4. Emotional Information Processing:
[1960] The server receives the user's emotional information sent from the terminal and adjusts the content and tone of the response to the user based on this information.
[1961] 5. Virtual environment display:
[1962] The server generates data to visually display each stage of the beer production process and sends it to the device, allowing users to experience it in a realistic way through visual content.
[1963] Terminal handling
[1964] 1. User Interface:
[1965] The terminal acts as a virtual tour guide, displaying each stage of the beer production process in the user's language of choice, with an intuitive and easy-to-use interface.
[1966] 2. Speech recognition and output:
[1967] The device converts the user's voice input into text and synthesizes the text received from the server into speech for the user to hear. The libraries used are "speech_recognition" and "pyttsx3".
[1968] 3. Emotion recognition:
[1969] The device recognizes the user's emotions using voice input and facial recognition technology, and sends the information to the server. Emotion recognition is performed using OpenCV and a pre-trained emotion recognition model.
[1970] 4. Data synchronization:
[1971] The device periodically synchronizes learning progress with the server, providing the user with the most up-to-date information.
[1972] User Behavior
[1973] 1. Initial Settings and Selections:
[1974] Users launch the app, select the language they want to learn, and enter basic information, which the device then sends to the server.
[1975] 2. Selection of beer production process:
[1976] Within the app, users select the stage of the beer-making process they want to learn about (e.g., fermentation), and the device sends this information to the server, which then generates a corresponding dialogue scenario.
[1977] 3. Execute the dialogue:
[1978] The user interacts with the system at each stage of the process to learn. For example, in response to a question about the fermentation process, "Tell me how alcohol is produced during fermentation," the server responds, "In the fermentation process, yeast converts sugars into alcohol and carbon dioxide."
[1979] 4. Emotion Recognition and Processing:
[1980] If the user shows an emotional reaction to a question (e.g., a questioning or surprised expression), the device recognizes that emotion and transmits it to the server. The server then adjusts the tone and content of the response based on the emotional information, and generates an appropriate response and transmits it to the device.
[1981] 5. Continuing the dialogue:
[1982] If the user asks a further question (e.g., "What types of yeast are there?"), the device recognizes the question and sends it to the server, which generates an appropriate response, which the device then speaks to the user.
[1983] Prompt Sentence Examples
[1984] "A Guide to the Beer Making Process: Tell Me About Fermentation."
[1985] "A Guide to the Beer Making Process: Explain the Malting Steps."
[1986] In this way, users can deepen their language learning and intercultural understanding through natural dialogue while learning about the beer brewing process. The introduction of an emotion engine provides responses that correspond to the user's emotions, creating a more personalized learning experience.
[1987] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1988] Step 1:
[1989] The server receives language selection information and basic information entered by the user through the device, including profile information such as the user's preferred language, name, and age. The received information is stored in a database and used to generate a learning plan.
[1990] Step 2:
[1991] The server generates a dialogue scenario about the beer brewing process. It uses a generative AI model to create a natural dialogue scenario based on the language selected by the user. This dialogue scenario contains content related to the stage of the process the user wants to learn about (e.g., fermentation). The input used here is the user information saved in step 1, and the output is the generated dialogue scenario.
[1992] Step 3:
[1993] The terminal provides the generated dialogue scenario in the language selected by the user. The user interface (UI) displays the dialogue scenario visually and audibly, allowing the user to begin the dialogue. The input here is the dialogue scenario sent from the server, and the output is the dialogue scenario displayed on the user interface.
[1994] Step 4:
[1995] The user inputs a question or response into the terminal. Speech recognition technology is used to convert the user's voice input into text, and the text is sent to the server. The input is the user's voice, and the output is the user's question or response converted into text.
[1996] Step 5:
[1997] The server recognizes the questions and responses entered by the user and generates an appropriate response. It uses a generative AI model to perform natural language processing and generate an appropriate response to the user's question. The input is the text received in step 4, and the output is the generated response.
[1998] Step 6:
[1999] The terminal outputs the response received from the server to the user as voice. It uses speech synthesis technology (pyttsx3) to convey information in a way that is easy for the user to understand. The input is the text of the response sent from the server, and the output is the voice response.
[2000] Step 7:
[2001] The device recognizes the user's emotions and sends the information to the server. Emotion recognition uses OpenCV and a pre-trained emotion recognition model. The input is the user's facial image and voice, and the output is the recognized emotion information.
[2002] Step 8:
[2003] The server adjusts the content and tone of the response based on the emotional information. It changes the tone and content of the response depending on the emotion expressed by the user, generating a more personalized response. The input is the emotional information received in step 7, and the output is the adjusted response.
[2004] Step 9:
[2005] The device outputs the adjusted response to the user as voice, again using speech synthesis technology to provide an adapted response to the user. The input is the adjusted response text sent from the server, and the output is the voice response.
[2006] Step 10:
[2007] The terminal visually displays the beer brewing process in a virtual environment. Using virtual reality and augmented reality technology, the user can visually experience the beer brewing process. The input is visual data sent from the server, and the output is the beer brewing process displayed in the virtual environment.
[2008] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[2009] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[2010] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[2011] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[2012] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[2013] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[2014] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[2015] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[2016] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[2017] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[2018] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[2019] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[2020] In the above embodiment, an example in which the ...
Claims
1. means for receiving language selection information and basic information input by a user; means for generating dialogue scenarios relating to a beer production process; means for providing the generated dialogue scenario in a language selected by the user; means for recognizing questions and responses entered by a user and generating appropriate responses; means for outputting the generated response to a user; A system including:
2. 10. The system of claim 1, further comprising means for recording the user's learning progress and transmitting the recording to the server.
3. 10. The system of claim 1, further comprising means for periodically adding to the database and providing to the user up-to-date information regarding the beer production process.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A