System

The system addresses the challenge of natural language acquisition by generating personalized study plans and conducting multilingual conversations, enabling effective bilingual development through integrated daily interactions.

JP2026028674APending Publication Date: 2026-02-20SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024131290
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-07
Publication Date
2026-02-20

AI Technical Summary

Technical Problem

Conventional language learning systems fail to provide an effective and natural way for individuals, especially children, to acquire a second native language, particularly in a multilingual environment, leading to difficulties in achieving fluent multilingual conversation skills or business-level language proficiency.

Method used

A system that includes inputting user information, generating a personalized study plan, conducting multilingual conversations using a natural language processing engine, evaluating progress, and adjusting the plan to develop listening, speaking, reading, and writing skills through a robotic device.

Benefits of technology

Creates an environment where multilingual conversations are integrated into daily life, allowing users to naturally acquire a second language and develop bilingual skills effectively, especially from an early age.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026028674000001_ABST
    Figure 2026028674000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for inputting information of a user and generating a learning plan based on the information; means for conducting a multilingual conversation using a natural language processing engine based on the generated learning plan; means for evaluating progress of the user and modifying or adjusting the learning plan based on the evaluation; and means for providing activities for practicing listening, speaking, reading, and writing skills based on the generated learning plan.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Conventional language learning systems make it difficult to learn a second native language in a natural and efficient way, posing particular challenges for language acquisition from an early age. In conventional education systems, it is difficult for children aiming to become bilingual to naturally come into contact with a multilingual environment, making it difficult for them to acquire fluent multilingual conversation skills or language proficiency at a business level. There is a need to address this issue. [Means for solving the problem]

[0005] The present invention solves the above problems by providing a system that includes: means for inputting user information and generating a study plan based on the information; means for conducting multilingual conversations using a natural language processing engine based on the generated study plan; means for evaluating the user's progress and modifying or adjusting the study plan based on the evaluation; and means for providing activities for practicing listening, speaking, reading, and writing skills based on the generated study plan.

[0006] Specifically, by registering user information in a database, creating a learning plan based on that information, and using a robotic device to converse with the user in real time, we provide a learning environment in which a second native language can be acquired in a natural way at home or in everyday life. As a result, we create an environment in which multilingual conversations are constantly taking place, and we realize that users can grow in a natural bilingual environment.

[0007] "User" refers to any individual who uses the system, including children and adults, particularly those who wish to learn a language.

[0008] "Information" refers to data about the user, including name, age, language level, learning goals, etc.

[0009] "Study Plan" means an individualized language learning plan generated based on user information, including a plan designed to develop listening, speaking, reading, and writing skills in a balanced manner.

[0010] "Generating means" refers to the act or process of creating a learning plan based on user information, including the software and algorithms involved.

[0011] "Natural language processing engine" refers to a program for understanding and generating human language, including those for smoothly conducting multilingual conversations.

[0012] "Multilingual conversation" refers to a dialogue that takes place in two or more languages, including real-time communication between a user and a system.

[0013] "Progress" refers to the results and accomplishments of a user's language learning in accordance with a learning plan, and includes data for evaluating this.

[0014] "Listening" refers to the practice of listening skills and includes activities designed to develop the ability to understand audio information.

[0015] "Speaking" refers to the practice of speaking skills and includes activities designed to develop the ability to communicate ideas and information using language.

[0016] "Reading" refers to the practice of reading skills and includes activities designed to develop the ability to understand sentences and texts.

[0017] "Writing" refers to the practice of writing skills and includes activities designed to develop the ability to produce sentences and texts.

[0018] The term "robot device" refers to a mechanical device for interacting with a user, and includes hardware and software for conducting natural conversation in real time. [Brief explanation of the drawings]

[0019] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5]FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0020] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0021] First, the terms used in the following description will be explained.

[0022] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0023] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0024] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0025] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0026] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0027] [First embodiment]

[0028] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0029] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0030] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0031] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0032] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0033] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0034] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0035] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0036] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0037] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0038] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0039] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0040] The system of the present invention allows users to input their information, generate an individually customized learning plan based on that information, and engage in multilingual conversations using a generation AI. It also evaluates progress and adjusts the learning plan based on that evaluation, enabling the development of listening, speaking, reading, and writing skills in a balanced manner.

[0041] 1. Entering user data and registering

[0042] The server receives new user information and registers it in a database. User data includes name, age, language level, learning goals, etc.

[0043] Examples:

[0044] A user accesses the application and enters information such as name, age, language level, learning goals, etc. on the new registration screen. This information is received by the server and registered in the database.

[0045] 2. Initial setup and customization

[0046] The device generates an individual learning plan based on the user's information, which is customized according to the user's age, language level, and learning goals.

[0047] Examples:

[0048] When a 5-year-old user is enrolled, the device generates a learning plan for young children, including games, songs, and activities using basic English words and phrases.

[0049] 3. Multilingual conversation function

[0050] The device sets a schedule based on the learning plan and starts a multilingual conversation session according to that schedule. The conversation engine uses natural language processing to converse with the user in real time.

[0051] Examples:

[0052] At 9 a.m., the device speaks to the user, saying, "It's English time. Good morning! How did you sleep last night?" This dialogue is facilitated by generative AI, which automatically generates the next question or comment based on the user's response.

[0053] 4. Four English Skills Approach

[0054] The device offers a variety of activities to balance and develop users' listening, speaking, reading and writing skills.

[0055] Examples:

[0056] During the listening section, the device will play a short story to the user and then ask questions about the story, while during the speaking section, the device will have the user pronounce simple phrases and evaluate their pronunciation to guide them to the next step.

[0057] 5. Mother Tongue Acquisition Program

[0058] The server periodically evaluates the user's learning progress and analyzes the progress data stored in the database, and based on this evaluation, modifies or adjusts the learning plan.

[0059] Examples:

[0060] At the end of the month, during progress assessment, the server determines that the user's listening skills have improved, but their speaking skills are lacking. Based on this, speaking-focused activities are added to the next month's lesson plan.

[0061] The system of the present invention naturally integrates multilingual exposure into the user's living environment, allowing the user to naturally acquire a second language in their daily life and grow as a bilingual. This system is particularly effective in language acquisition from an early age, and solves the problems facing the modern education system.

[0062] The processing flow will be explained below.

[0063] Processing flow

[0064] Entering user data and registering

[0065] Step 1:

[0066] Users access the application and enter information such as their name, age, language level, and learning goals on the new registration screen.

[0067] Step 2:

[0068] The terminal receives the entered user information and sends it to the server.

[0069] Step 3:

[0070] The server registers the received user information in a database and generates a user ID.

[0071] First-time setup and customization

[0072] Step 1:

[0073] The server retrieves user information from a database and generates an initial learning plan based on this.

[0074] Step 2:

[0075] The device presents the generated learning plan to the user and asks for confirmation of the learning plan.

[0076] Step 3:

[0077] The user reviews the proposed study plan and requests adjustments if necessary.

[0078] Multilingual conversation function

[0079] Step 1:

[0080] The device starts a conversation session on a specified schedule based on the learning plan.

[0081] Step 2:

[0082] MimiTomo (robot) uses a natural language processing engine to converse with users in the selected language.

[0083] Step 3:

[0084] Users have the opportunity to converse with the robot and use language naturally.

[0085] Four English Skills Approach

[0086] Step 1:

[0087] The device generates and provides activities for practicing listening, speaking, reading and writing skills.

[0088] Step 2:

[0089] The user performs each skill practice activity provided by the terminal.

[0090] Step 3:

[0091] The terminal collects the user's activity results and sends them to the server.

[0092] Mother Tongue Acquisition Program

[0093] Step 1:

[0094] The server periodically evaluates the user's learning progress and updates the database based on the evaluation results.

[0095] Step 2:

[0096] The server analyzes the progress data and adjusts the learning plan as needed.

[0097] Step 3:

[0098] The device will present the user with a new, adjusted learning plan and begin the next learning cycle.

[0099] In this way, the system generates an individualized learning plan based on the user's information, engages in daily multilingual conversations based on that plan, and provides activities to train the listening, speaking, reading, and writing skills in a balanced manner.By continuously evaluating the user's progress, the system supports effective bilingual development.

[0100] Example 1

[0101] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0102] In multilingual learning, there is a need for systems that can take into account individual differences, provide users with a learning plan that is tailored to their needs, and adjust it as needed based on their progress. In particular, there is a lack of effective methods for balanced development of listening, speaking, reading, and writing skills. Another issue is the difficulty of effectively incorporating real-time multilingual conversation functions.

[0103] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0104] In this invention, the server includes means for inputting user information and generating a study plan based on the information, means for conducting multilingual conversations using a natural language processing engine based on the generated study plan, means for evaluating the user's progress and changing or adjusting the study plan based on the evaluation, and means for providing activities for practicing the skills of auditory comprehension, speaking, reading comprehension, and writing based on the generated study plan. This makes it possible to create and adjust a study plan suited to the user and develop each skill in a balanced manner.

[0105] "User" refers to an individual who uses the system to acquire multiple languages.

[0106] "Information" refers to data about the user such as name, age, language level, learning goals, etc.

[0107] "Study Plan" refers to a customized study schedule and content based on user information.

[0108] A "natural language processing engine" refers to technology that enables computers to understand and generate human language.

[0109] A "multilingual conversation" refers to a dialogue that takes place using more than one language.

[0110] "Progress" refers to the state of how far a user has progressed or improved on their learning plan.

[0111] "Evaluation" refers to the process of measuring a user's learning progress and analyzing the results.

[0112] "Auditory comprehension" refers to the skill of understanding language heard.

[0113] "Speech" refers to the skill of speaking a language.

[0114] "Reading comprehension" refers to the skill of understanding written language.

[0115] "Writing" refers to the skill of using language to write sentences.

[0116] "Activity" refers to specific exercises or tasks that a user performs to practice each skill.

[0117] The system of the present invention inputs user information, generates an individually customized learning plan based on that information, and conducts multilingual conversations using a generative AI model. It also evaluates progress and adjusts the learning plan based on that evaluation, enabling the development of auditory comprehension, speaking, reading comprehension, and writing skills in a balanced manner.

[0118] 1. Entering user data and registering

[0119] The server receives new user information and registers it in a database. User data includes name, age, language level, learning goals, etc.

[0120] Examples:

[0121] A user accesses the application and enters information such as name, age, language level, learning goals, etc. on the new registration screen. This information is received by the server and registered in the database.

[0122] 2. Initial setup and customization

[0123] The device generates an individual learning plan based on the user's information, which is customized according to the user's age, language level, and learning goals.

[0124] Examples:

[0125] If a 5-year-old user is enrolled, the device will generate a learning plan for young children, including games, songs, and activities using basic English words and phrases.

[0126] 3. Multilingual conversation function

[0127] The device sets a schedule based on the study plan and starts a multilingual conversation session according to that schedule. The conversation engine uses natural language processing to converse with the user in real time.

[0128] Examples:

[0129] At 9 a.m., the device speaks to the user, "It's time for English. Good morning! How did you sleep last night?" This dialogue is driven by a generative AI model, which automatically generates the next question or comment based on the user's response.

[0130] 4. Four English Skills Approach

[0131] The device offers a variety of activities to balance the user's auditory comprehension, speaking, reading, and writing skills.

[0132] Examples:

[0133] During the listening section, the device will play a short story to the user and then ask questions about the story. During the speaking section, the device will have the user pronounce simple phrases and evaluate their pronunciation to guide them to the next step.

[0134] 5. Mother Tongue Acquisition Program

[0135] The server periodically evaluates the user's learning progress and analyzes the progress data stored in the database, and based on this evaluation, modifies or adjusts the learning plan.

[0136] Examples:

[0137] At the end of the month, the server assesses the progress and finds that the user's listening skills have improved, but their speaking skills are lacking. Based on this, activities to strengthen speaking are added to the next month's lesson plan.

[0138] Through this entire system, multilingual exposure can be naturally incorporated into the user's living environment, allowing them to naturally acquire a second language in their daily lives and grow as bilinguals. This is particularly effective for language acquisition from an early age, and solves the problems facing the modern education system.

[0139] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0140] Step 1:

[0141] Entering user data and registering

[0142] Input: The user enters information such as name, age, language level, and learning goals on the application's sign-up screen.

[0143] Specific behavior: The data entered by the user is sent to the server through the input form.

[0144] Data processing: The server formats the received user data into an appropriate format (e.g., converts it to JSON format).

[0145] Output: Correctly formatted user data.

[0146] Step 2:

[0147] Receiving information and registering in the database

[0148] Input: The server receives the formatted user data.

[0149] What happens: The server checks the format of the data and issues the appropriate SQL query to store it in the database.

[0150] Data manipulation: Insert user data into the database using SQL queries.

[0151] Output: User data is registered in the database.

[0152] Step 3:

[0153] User information acquisition

[0154] Input: User data registered on the server.

[0155] Specific operation: The device obtains user data from the server through an API call.

[0156] Data processing: Converting received data into internal data structures.

[0157] Output: User data stored in the device.

[0158] Step 4:

[0159] Generate a lesson plan

[0160] Input: User data stored on the device.

[0161] What it does: The lesson plan generator generates an appropriate lesson plan based on the user's age, language level, and learning goals.

[0162] Data processing: The study plan generator analyzes user data and creates a customized study plan that is optimal for them.

[0163] Output: A customized learning plan for the user.

[0164] Step 5:

[0165] Scheduling

[0166] Input: Schedule information based on your study plan.

[0167] Specific behavior: The device uses the calendar API to automatically reserve a session for the specified date and time.

[0168] Data processing: Convert schedule information into calendar format and register it on the user's device.

[0169] Output: The configured schedule.

[0170] Step 6:

[0171] Starting a multilingual conversation session

[0172] Input: A prompt statement based on the lesson plan.

[0173] Specific operation: The device sends a prompt to the generative AI model and begins a dialogue with the user in real time.

[0174] Data processing: Input the prompt sentence into the natural language processing engine and receive a response from the generative AI model.

[0175] Output: An ongoing multilingual conversation with the user.

[0176] Step 7:

[0177] Skills Assessment

[0178] Input: User behavior data during a multilingual conversation session.

[0179] How it works: The device uses speech recognition technology and natural language processing to evaluate the user's pronunciation and responses.

[0180] Data processing: Analyze the collected data and generate metrics to assess the user's skill level.

[0181] Output: User skill assessment results.

[0182] Step 8:

[0183] Collecting and evaluating progress data

[0184] Input: Skill assessment results after each session.

[0185] Specific operation: The server collects and analyzes the user's learning progress data sent from the device.

[0186] Data processing: The collected data is input into a statistical model to evaluate progress.

[0187] Output: Progress assessment results and reports.

[0188] Step 9:

[0189] Adjusting your study plan

[0190] Input: Progress assessment results.

[0191] Specific operation: The server generates a new learning plan based on the progress assessment and distributes it to the device.

[0192] Data processing: Applying algorithms to progress data to create new, customized learning plans.

[0193] Output: A tailored learning plan.

[0194] (Application example 1)

[0195] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0196] Increasing the efficiency of multilingual learning is a challenge facing the modern education system. Furthermore, with limited options for effective language learning in busy daily lives, there is a need for learning methods that allow for effective use of time, especially while traveling. There is also a need for learning plans tailored to the age and language level of students, from children to adults, that foster a balanced development of listening, speaking, reading, and writing skills. To address these challenges, a system is needed that can provide customized learning plans and progress management tailored to each user's individual needs.

[0197] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0198] The invention includes means for inputting user information and generating a learning plan based on the information, means for conducting multilingual conversations using a natural language processing engine based on the generated learning plan, means for evaluating the user's progress and changing or adjusting the learning plan based on the evaluation, means for providing activities to practice listening, speaking, reading, and writing skills based on the generated learning plan, and means for providing multilingual learning using a display, microphone, and speaker in the vehicle.

[0199] This allows passengers to effectively utilize their time while traveling in a vehicle to study multiple languages, and enables them to efficiently develop their language skills in a balanced manner through a learning plan customized for each user.In addition, real-time dialogue and progress management can maximize the effectiveness of their learning.

[0200] "User Information" refers to personal data necessary to generate a learning plan, such as the learner's name, age, language level, and learning goals.

[0201] "Study Plan" refers to specific learning content and schedules that are individually customized based on the user's information and designed to develop listening, speaking, reading, and writing skills in a balanced manner.

[0202] A "natural language processing engine" refers to a software system that uses generative AI to conduct natural conversations in multiple languages.

[0203] "Progress assessment" refers to the process of quantitatively and qualitatively evaluating the results of the learning activities that a user has carried out based on their learning plan.

[0204] "Listening, speaking, reading and writing skills" refers to the four basic skill areas in language learning, with corresponding practice activities.

[0205] "In-vehicle display" refers to the screen installed inside an autonomous vehicle that displays learning content and provides interactive instructions.

[0206] A "microphone" refers to an input device that picks up a user's voice and analyzes the voice in real time.

[0207] "Speaker" refers to an audio device that allows interaction between the user and the system through audio output.

[0208] A "robot device" refers to a hardware device used to interact with a user in real time, and has functions such as voice recognition and voice synthesis.

[0209] This invention is a system for providing multilingual learning within an autonomous vehicle, which has the function of generating a learning plan based on user information and conducting multilingual conversations in real time based on that plan. The system includes the following elements:

[0210] 1. Entering user data and registering

[0211] The server receives information such as the user's name, age, language level, and learning goals provided by the user via a display, smartphone, or other device when they board an autonomous vehicle, and registers this information in a database, creating a foundation for generating learning plans tailored to the user's individual needs.

[0212] 2. Create and customize your study plan

[0213] The server uses a generative AI model based on registered user information to generate individual learning plans. For example, it can provide a business English course for adults commuting to work, and a playful basic English learning plan for children. The learning plans include a balanced mix of listening, speaking, reading, and writing skills.

[0214] 3. Multilingual conversation function

[0215] The device sets a schedule based on the learning plan and starts multilingual conversation sessions according to that schedule. Using the vehicle's microphone and speaker, the device can communicate with the user in real time using a natural language processing engine, allowing users to efficiently improve their language skills while on the move.

[0216] 4. Four Skills Approach

[0217] The device offers multifaceted activities to practice listening, speaking, reading, and writing skills based on a learning plan. For example, during listening, the device plays a short story followed by questions to check comprehension. During speaking, the device evaluates the user's pronunciation in real time and provides feedback.

[0218] 5. Progress assessment and plan adjustment

[0219] The server periodically evaluates the user's progress and analyzes the progress data stored in the database. Based on this evaluation, the server changes or adjusts the learning plan to help the user improve their skills. For example, if listening skills are improving but speaking skills are lacking, the server will add activities to strengthen speaking in the next month.

[0220] Specific examples

[0221] A user boards an autonomous vehicle and registers by entering their name, age, language level, and learning goals on a display terminal. The server then generates a customized learning plan based on the user's information and initiates a multilingual conversation session through the vehicle's speakers and display.

[0222] Prompt Sentence Examples

[0223] "Enter the user's name, age, language level, and learning goals, and generate a program to charge them in the car's interface."

[0224] This system allows users to study multiple languages ​​according to an individually customized learning plan, making effective use of their time inside an autonomous vehicle. It is expected that real-time dialogue and progress management will maximize learning effectiveness.

[0225] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0226] Step 1:

[0227] When a user gets into an autonomous vehicle, they input information via a display or smartphone. This information includes their name, age, language level, and learning goals, and the input is sent to a server. The server receives this information and registers it in a database. The input of this step is user information, and the output is the user information registered in the database.

[0228] Step 2:

[0229] The server uses a generative AI model to generate individual study plans based on registered user information. The generated study plans are customized according to the user's age, language level, and learning goals. For example, an adult user aiming for business English would be provided with a study plan focused on business situations. The input for this step is user information registered in the database, and the output is a customized study plan.

[0230] Step 3:

[0231] The device sets a schedule based on the learning plan sent from the server. Specifically, it uses the vehicle's display and speaker to indicate the required activities to the user and start the learning session. For example, it sets an English listening session to start at 9:00 AM. The input of this step is the customized learning plan, and the output is the set schedule.

[0232] Step 4:

[0233] The device initiates a multilingual conversation session according to a set schedule. It collects user utterances through a microphone in the vehicle and analyzes them through a natural language processing engine. It uses a generative AI model on the analyzed data to generate an appropriate reply and respond to the user through the speaker. The input for this step is the user's utterance data, and the output is the generated response.

[0234] Step 5:

[0235] The device provides activities for practicing listening, speaking, reading, and writing skills. For example, a listening session plays a short story followed by questions to check comprehension. A speaking session allows the user to practice pronunciation and evaluate its accuracy with a speech recognition system. The input for this step is the user's activity data based on the learning plan, and the output is the evaluation result.

[0236] Step 6:

[0237] The server periodically evaluates the user's progress and analyzes the progress data stored in the database. Based on this evaluation, it modifies or adjusts the user's learning plan. For example, if the user's listening skills are improving but their speaking skills are lacking, it will add activities to strengthen their speaking skills next month. The input of this step is the progress data stored in the database, and the output is an adjusted learning plan.

[0238] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0239] The system generates a learning plan based on user information, and utilizes a natural language processing engine and an emotion engine to evaluate the user's progress and emotions while conducting multilingual conversations, dynamically adjusting the learning plan. By nurturing listening, speaking, reading, and writing skills in a balanced manner and taking into account the user's emotional state, the system provides an optimal learning experience.

[0240] 1. Entering user data and registering

[0241] The server receives new user information and registers it in a database. User data includes name, age, language level, learning goals, etc.

[0242] Examples:

[0243] A user accesses the application and enters information such as name, age, language level, learning goals, etc. on the new registration screen. This information is received by the server and registered in the database.

[0244] 2. Initial setup and customization

[0245] The device generates an individual learning plan based on the user's information, which is customized according to the user's age, language level, and learning goals. Furthermore, an emotion engine is used to configure the plan to take into account the user's emotional state.

[0246] Examples:

[0247] When a five-year-old user registers, the device generates a learning plan for young children, including games and songs, activities using basic English words and phrases, and selects content that will interest the user and analyzes it with an emotion engine.

[0248] 3. Multilingual conversation function

[0249] The device sets a schedule based on the learning plan and starts a multilingual conversation session according to that schedule. The conversation engine uses natural language processing to converse with the user in real time. Furthermore, the emotion engine analyzes the user's emotions and adjusts the content of the conversation accordingly.

[0250] Examples:

[0251] At 9 a.m., the device speaks to the user, saying, "It's time for English. Good morning! How did you sleep last night?" This dialogue is conducted by a generative AI, which automatically generates the next question or comment based on the user's response. If the user shows no interest, the emotion engine analyzes the user's facial expressions and tone of voice and changes the topic.

[0252] 4. Four English Skills Approach

[0253] The device generates and provides activities to balance listening, speaking, reading, and writing skills, and uses an emotion engine to analyze the user's emotional state and provide appropriate feedback.

[0254] Examples:

[0255] During the listening section, the device tells the user a short story and then asks them questions about it. During the speaking section, the device asks the user to pronounce simple phrases, and the device evaluates their pronunciation and guides them to the next step. An emotion engine analyzes whether the user is enjoying themselves or struggling and adjusts its approach accordingly.

[0256] 5. Mother Tongue Acquisition Program

[0257] The server periodically evaluates the user's learning progress and analyzes the progress data stored in the database. Based on this evaluation, the server changes or adjusts the learning plan. It also uses an emotion engine to consider the user's emotional state and provide an optimal plan that maintains the user's motivation to learn.

[0258] Examples:

[0259] At the end of the month, the server will assess the user's progress and determine that their listening skills have improved, but their speaking skills are lacking. Based on this, activities to strengthen speaking will be added to the next month's learning plan. Furthermore, the emotion engine will analyze the user's motivation and optimize the learning plan to increase activities that they find challenging.

[0260] This system allows users to naturally acquire a second language in their daily lives and grow as bilinguals. By combining it with an emotion engine, it is possible to grasp the user's emotional state in real time and provide an optimal learning environment. As a result, an environment is created where multilingual conversations are constantly taking place, allowing users to grow naturally in a bilingual environment.

[0261] The processing flow will be explained below.

[0262] Processing flow

[0263] Entering user data and registering

[0264] Step 1:

[0265] Users access the application and enter information such as their name, age, language level, and learning goals on the new registration screen.

[0266] Step 2:

[0267] The terminal receives the entered user information and sends it to the server.

[0268] Step 3:

[0269] The server registers the received user information in a database and generates a user ID.

[0270] First-time setup and customization

[0271] Step 1:

[0272] The server retrieves user information from a database and generates an initial learning plan based on this.

[0273] Step 2:

[0274] The device presents the generated learning plan to the user and asks for confirmation of the learning plan.

[0275] Step 3:

[0276] The user reviews the proposed study plan and requests adjustments if necessary.

[0277] Step 4:

[0278] The server adjusts the learning plan based on user feedback and determines the final plan.

[0279] Multilingual conversation function

[0280] Step 1:

[0281] The device starts a conversation session on a specified schedule based on the learning plan.

[0282] Step 2:

[0283] MimiTomo (robot) uses a natural language processing engine to converse with users in the selected language.

[0284] Step 3:

[0285] The emotion engine analyzes the user's facial expressions and tone of voice to determine the user's emotional state.

[0286] Step 4:

[0287] MimiTomo (the robot) adjusts the conversation based on the user's emotional state, for example, changing the topic if the user is losing interest.

[0288] Step 5:

[0289] Users have the opportunity to converse with the robot and use language naturally.

[0290] Four English Skills Approach

[0291] Step 1:

[0292] The device generates and provides activities for practicing listening, speaking, reading and writing skills.

[0293] Step 2:

[0294] The user performs each skill practice activity provided by the terminal.

[0295] Step 3:

[0296] The emotion engine analyzes the user's emotional state during an activity and adjusts the difficulty and content of the activity if motivation is declining.

[0297] Step 4:

[0298] The terminal collects the user's activity results and sends them to the server.

[0299] Mother Tongue Acquisition Program

[0300] Step 1:

[0301] The server periodically evaluates the user's learning progress and updates the database based on the evaluation results.

[0302] Step 2:

[0303] The server analyzes the progress data and adjusts the learning plan as needed.

[0304] Step 3:

[0305] The emotion engine also evaluates the user's emotional data and reflects this in adjusting the learning plan.

[0306] Step 4:

[0307] The device will present the user with a new, adjusted learning plan and begin the next learning cycle.

[0308] Specific examples

[0309] Entering user data and registering

[0310] Step 1:

[0311] A user accesses the application and registers by entering their name "Taro," age "5 years old," language level "Beginner," and learning goal "Bilingual."

[0312] Step 2:

[0313] The terminal receives the input information and transmits it to the server.

[0314] Step 3:

[0315] The server registers the received information in the database and generates a user ID "001."

[0316] First-time setup and customization

[0317] Step 1:

[0318] The server generates an individual learning plan, "English Learning Plan for Toddlers," based on the user ID "001."

[0319] Step 2:

[0320] The device will present the generated learning plan to the user and ask them to "confirm."

[0321] Step 3:

[0322] The user reviews the plan and requests any necessary adjustments (e.g., adding specific activities).

[0323] Step 4:

[0324] The server adjusts the learning plan based on the request and finalizes the plan.

[0325] Multilingual conversation function

[0326] Step 1:

[0327] The device will set up and start an English conversation session at 9am every morning.

[0328] Step 2:

[0329] MimiTomo (robot) says, "Good morning, Taro! How did you sleep last night?"

[0330] Step 3:

[0331] The emotion engine analyzes the user's facial expression and determines that they are "showing interest."

[0332] Step 4:

[0333] MimiTomo (robot) continues the topic by saying, "Let's talk about your favorite toy."

[0334] Step 5:

[0335] The user responds in English, "My favorite toy is a car."

[0336] Four English Skills Approach

[0337] Step 1:

[0338] The device generates a listening activity and plays a short story.

[0339] Step 2:

[0340] Users listen to a story and answer related questions.

[0341] Step 3:

[0342] The emotion engine analyzes the user's level of concentration and determines that they are "concentrated."

[0343] Step 4:

[0344] The device provides additional activity questions to promote deeper understanding.

[0345] Mother Tongue Acquisition Program

[0346] Step 1:

[0347] The server evaluates the user's progress at the end of each month and updates the database.

[0348] Step 2:

[0349] The server analyzes the progress data and determines that "speaking skills need to be improved."

[0350] Step 3:

[0351] The emotion engine analyzes the user's motivation data and selects activities that they find rewarding.

[0352] Step 4:

[0353] The device will present the user with a new, adjusted learning plan and begin the next learning cycle.

[0354] Example 2

[0355] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0356] Conventional educational systems have had difficulty dynamically adjusting learning plans based on each learner's progress and emotional state. Especially in multilingual conversation learning, there has been a lack of appropriate feedback and activities that take into account the learner's emotional state. As a result, it has been difficult to maintain learner motivation and acquire skills efficiently.

[0357] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0358] In this invention, the server includes means for inputting user information and generating a study plan based on the information, means for conducting multilingual conversations using a natural language processing engine based on the generated study plan, means for evaluating the user's progress and changing or adjusting the study plan based on the evaluation, means for training the listening, speaking, reading, and writing skills in a balanced manner, and means for analyzing the user's emotional state and dynamically adjusting the study plan, thereby making it possible to provide an optimal learning experience according to the progress and emotional state of each individual learner.

[0359] "Means for inputting user information" refers to a method or device by which a user inputs information such as name, age, language level, learning goals, etc. into the system.

[0360] A "means for generating a study plan" is a method or device for creating an individualized study plan based on user information.

[0361] A "natural language processing engine" is a technology that uses artificial intelligence to analyze user input and generate appropriate responses.

[0362] A "means for conducting a multilingual conversation" is a method or device for conducting a conversation with a user in multiple languages.

[0363] A "means for assessing user progress" is a method or device for measuring and analyzing a user's learning progress.

[0364] A "means for changing or adjusting a study plan" is a method or device for modifying an existing study plan based on a user's progress or needs.

[0365] "A means for balanced training of listening, speaking, reading, and writing skills" is a method or device for developing these language skills in a balanced manner.

[0366] The "means for analyzing the user's emotional state" refers to a method or device for analyzing the user's facial expression, tone of voice, etc. to determine their emotions.

[0367] A "means for dynamically adjusting a study plan" is a method or apparatus for optimizing a study plan in real time according to a user's emotional state and progress.

[0368] An "information processing device" is a computer system that performs multilingual conversations, assessment of learning progress, analysis of emotional states, and so on.

[0369] This invention relates to a system that generates a study plan based on user information, and dynamically adjusts the study plan by evaluating the user's progress and emotions while conducting multilingual conversations using a natural language processing engine and an emotion engine. This system is configured to effectively support the user's language learning.

[0370] The system mainly consists of a server and a terminal. The server receives user information and registers it in a database. User data includes name, age, language level, learning goals, etc. The server accesses the database and generates an individualized initial learning plan based on the user information. The terminal receives this learning plan and configures it to predict the user's emotional state. At this point, an emotional engine (e.g., Affectiva) is used to analyze the user's emotional state.

[0371] The learning plan is designed to balance the training of listening, speaking, reading, and writing skills. The device will speak to the user according to a set schedule, for example at 9:00 AM, asking, "Good morning! How did you sleep last night?" This dialogue is conducted using a natural language processing engine (e.g., GPT-3), which generates an appropriate response based on the user's input. If the user shows no interest, the emotion engine analyzes the user's facial expressions and tone of voice and changes the topic.

[0372] The device also offers listening, speaking, reading, and writing activities. For example, during listening, it plays a short story followed by questions about the story. When the user answers the questions, the device evaluates the answer and guides the user to the next step. The emotion engine analyzes the user's reaction and provides appropriate feedback.

[0373] The server periodically retrieves the user's learning progress data from the database and evaluates that progress. For example, if listening skills are improving but speaking skills are lacking, activities to strengthen speaking will be added to the next month's learning plan. The emotion engine also analyzes the user's motivation and optimizes the plan to increase activities that are likely to be rewarding. The server then sends the updated learning plan to the device, allowing for a smooth start to the next month's learning.

[0374] This system allows users to learn multiple languages ​​naturally in their daily lives, and by combining it with an emotion engine, it is possible to grasp the user's emotional state in real time and provide the optimal learning environment. Multilingual conversations utilizing generative AI models and prompt sentences allow users to constantly have new learning experiences. By utilizing technologies such as "generative AI models and prompt sentences," the system provides the most effective learning method for users.

[0375] Examples:

[0376] Example prompt: "Good morning! How did you sleep last night?" The next question or comment is automatically generated based on the user's response.

[0377] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0378] System program processing steps

[0379] Step 1: Entering user data and registering

[0380] 1-1: The user opens the new registration screen.

[0381] Input: The user enters information such as name, age, language level, and learning goals.

[0382] 1-2: The server receives the information sent by the user.

[0383] Data processing: Convert the received user information into database format.

[0384] 1-3: The server registers the information in the database.

[0385] Output: Generates a message confirming successful registration and sends it to the terminal.

[0386] What it does: The server uses a database management system (e.g., MySQL) to store data.

[0387] Step 2: Initial setup and customization

[0388] 2-1: The server reads the user information in the database.

[0389] Input: User information retrieved from the database.

[0390] 2-2: The server analyzes the user information and generates an initial learning plan.

[0391] Data processing: Generate a customized learning plan based on the user's age, language level, and learning goals.

[0392] 2-3: The device is configured to predict the user's emotional state based on the learning plan received from the server.

[0393] Input: The learning plan sent by the server.

[0394] 2-4: The device presents the learning plan to the user and asks for confirmation.

[0395] Output: Generates a preview of the learning plan and displays it to the user.

[0396] Specific operation: The device configures the emotion engine (e.g., Affectiva) and prepares it to analyze the user's emotional state.

[0397] Step 3: Multilingual conversation function

[0398] 3-1: The device starts a conversation session according to the set schedule.

[0399] Input: Schedule information based on your study plan.

[0400] 3-2: The user talks to the terminal.

[0401] Input: User voice or text input.

[0402] 3-3: The terminal analyzes the user's input using a natural language processing engine.

[0403] Data Computing: Using a natural language processing engine (e.g., GPT-3) to analyze user input and generate appropriate responses.

[0404] 3-4: The device uses an emotion engine to analyze the user's emotional state and dynamically adjust the content of the conversation.

[0405] Input: User's facial expressions and tone of voice.

[0406] Output: Generates tailored conversation content and presents it to the user.

[0407] Specific operation: The device responds to the user with the conversation content generated by the device via voice or text.

[0408] Step 4: Four English Skills Approach

[0409] 4-1: The device starts the learning activity at the set time.

[0410] Input: Activity schedule based on the lesson plan.

[0411] 4-2: The devices provide listening, speaking, reading, and writing activities.

[0412] Input: User actions or responses.

[0413] Data processing: Record and analyze user responses to each activity.

[0414] 4-3: The device uses an emotion engine to evaluate the user's reaction and provide feedback as needed.

[0415] Input: The user's emotional state.

[0416] Output: Generate feedback based on the evaluation results.

[0417] What it does: The listening activity plays a short story, and the speaking activity teaches the user pronunciation.

[0418] Step 5: Mother Tongue Acquisition Program

[0419] 5-1: The server periodically retrieves the user's learning progress data from the database.

[0420] Input: Progress data accumulated in the database.

[0421] 5-2: The server analyzes the progress data and updates the learning plan.

[0422] Data calculations: Analyze progress data to identify skills that users need to improve.

[0423] 5-3: The server uses the emotion engine to evaluate the user's emotional state and optimize the learning plan.

[0424] Input: User sentiment analysis data.

[0425] Output: Generates an optimized learning plan and sends it to the device.

[0426] 5-4: The server sends the updated learning plan to the device.

[0427] Output: Sends the updated study plan to the device, ready to start studying for the next month.

[0428] What it does: The server uses analytics tools to evaluate progress data and generate a learning plan with appropriate changes.

[0429] (Application example 2)

[0430] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0431] Conventional multilingual conversation systems lack the ability to dynamically analyze and respond to users' progress and emotional state in real time, resulting in insufficient learning efficiency and user interaction. Furthermore, they lacked the information processing capabilities to provide personalized responses when dealing with customers in physical stores. This can result in a lack of motivation to learn and a decline in user satisfaction. In particular, flexible responses that take into account the user's emotional state are essential for complex multilingual interactions.

[0432] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0433] In this invention, the server includes: means for inputting user information and generating a study plan based on the information; means for conducting multilingual conversations using a natural language processing engine based on the generated study plan; means for evaluating the user's progress and changing or adjusting the study plan based on the evaluation; means for providing activities for practicing listening, speaking, reading, and writing skills based on the generated study plan; means including an emotion analysis engine for analyzing the user's emotional state and providing an optimal response in real time; and means including an information processing device for responding to customers in a store. This makes it possible to analyze the user's study progress and emotional state in real time and provide an optimal study plan based on the analysis, while also enabling personalized and flexible customer service in physical stores.

[0434] "User" refers to an individual who uses the system.

[0435] "Information" refers to data about the user such as name, age, language level, learning goals, etc.

[0436] "Study Plan" refers to a schedule and content of study activities customized based on user information.

[0437] A "natural language processing engine" refers to a software component that analyzes text data and conducts conversations in multiple languages.

[0438] "Multilingual conversation" refers to interaction using multiple languages, during which real-time translation and dialogue control are carried out.

[0439] "Progress" refers to the results and achievement of a user's learning activities.

[0440] "Listening" refers to the skill of understanding spoken information.

[0441] "Speaking" refers to the skill of conveying information orally.

[0442] "Reading" refers to the skill of reading and understanding written information.

[0443] "Writing" refers to the skill of describing information using written text.

[0444] "Emotional state" refers to a user's current mental and emotional state.

[0445] An "emotion analysis engine" refers to a software component that analyzes a user's emotional state and responds appropriately.

[0446] "Information processing device" refers to a device that includes hardware and software for inputting, processing, and outputting data.

[0447] "Store" means a physical location for conducting sales activities.

[0448] This system starts by inputting user information and generating a learning plan based on that information. Specifically, the server receives information such as the user's name, age, language level, and learning goals, and registers it in a pre-prepared database. This forms the basis for customizing an appropriate learning plan for each user.

[0449] Next, the device uses a natural language processing engine to conduct a multilingual conversation based on the generated learning plan. Specifically, a schedule is created based on the information entered by the user, and the conversation proceeds according to that schedule. During the conversation, an emotion analysis engine analyzes the user's emotional state in real time, and the content of the conversation is adjusted accordingly.

[0450] The system constantly evaluates the user's progress and changes or adjusts the learning plan based on the evaluation results. This method makes it possible to provide an optimal learning plan that is always up to date. For example, if the user's listening skills improve, new skill practice will be incorporated to advance to the next level.

[0451] Furthermore, the system provides activities to practice listening, speaking, reading, and writing skills in a balanced manner. During the activities, the system analyzes the user's emotional state and provides appropriate feedback. This analysis is performed using software components such as EmotionAnalyzer and NLPModel.

[0452] A specific example is an information processing device installed in a cash register that interacts with the user in real time. For example, a smartphone or information robot can provide personalized guidance based on the user's name and learning goals. If the user appears unhappy, an emotion analysis engine can analyze their emotional state and flexibly adapt the response to improve user satisfaction.

[0453] An example of a prompt is as follows:

[0454] "Customer ID: 1

[0455] Customer Name: Sato

[0456] Language: Japanese

[0457] Age: 30

[0458] Emotional state: happy

[0459] Response: Personalized greetings and product recommendations

[0460] Output: 'Hello, Sato-san. How can I help you today? We have great deals on electronics today.'

[0461] Based on this prompt, the generative AI model generates appropriate answers and responses, making customer service in physical stores much smoother.

[0462] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0463] Step 1:

[0464] The server inputs user information and registers it in a database. The input information includes name, age, language level, learning goals, etc. Specifically, when a user inputs information using a smartphone or interface device, the server receives it and stores it in a database.

[0465] Input: User's name, age, language level, learning goal

[0466] Output: User information registered in the database

[0467] Step 2:

[0468] The server generates a learning plan based on the information stored in the database, including content tailored to the user's age, language level, and learning goals, and uses a sentiment analysis engine to customize the plan based on the user's emotional state.

[0469] Input: User information registered in the database

[0470] Output: A customized study plan

[0471] Step 3:

[0472] The device initiates a multilingual conversation using a natural language processing engine based on the generated learning plan. Specifically, the device proceeds with a dialogue with the user according to the learning plan and generates appropriate responses in real time using a generative AI model.

[0473] Input: Customized Study Plan

[0474] Output: User interaction, real-time generated responses

[0475] Step 4:

[0476] The server periodically evaluates the user's progress and modifies or adjusts the learning plan based on the evaluation, which includes activity log data and user feedback.

[0477] Input: Learning activity log data, user feedback

[0478] Output: Adjusted learning plan

[0479] Step 5:

[0480] The device performs activities to practice listening, speaking, reading and writing skills based on a tailored learning plan, and during each activity, an emotion analysis engine analyzes the user's emotional state and provides appropriate feedback.

[0481] Input: Adjusted Study Plan

[0482] Output: User's emotional state data, feedback on activities

[0483] Step 6:

[0484] The server responds to users using the in-store information processing device. Specifically, when a user makes an inquiry to the in-store information processing device, the emotion analysis engine analyzes the user's emotional state based on the information and generates an appropriate response. Based on example prompt sentences, the generative AI model provides an appropriate response to the inquiry.

[0485] Input: In-store inquiry details, user's emotional state

[0486] Output: Personalized query response

[0487] Through the above processing steps, the system analyzes the user's learning progress and emotional state in real time, providing an optimal learning plan, and enabling personalized and flexible customer service in physical stores.

[0488] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0489] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0490] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0491] [Second embodiment]

[0492] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0493] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0494] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0495] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0496] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0497] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0498] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0499] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0500] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0501] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0502] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0503] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0504] The system of the present invention allows users to input their information, generate an individually customized learning plan based on that information, and engage in multilingual conversations using a generation AI. It also evaluates progress and adjusts the learning plan based on that evaluation, enabling the development of listening, speaking, reading, and writing skills in a balanced manner.

[0505] 1. Entering user data and registering

[0506] The server receives new user information and registers it in a database. User data includes name, age, language level, learning goals, etc.

[0507] Examples:

[0508] A user accesses the application and enters information such as name, age, language level, learning goals, etc. on the new registration screen. This information is received by the server and registered in the database.

[0509] 2. Initial setup and customization

[0510] The device generates an individual learning plan based on the user's information, which is customized according to the user's age, language level, and learning goals.

[0511] Examples:

[0512] When a 5-year-old user is enrolled, the device generates a learning plan for young children, including games, songs, and activities using basic English words and phrases.

[0513] 3. Multilingual conversation function

[0514] The device sets a schedule based on the learning plan and starts a multilingual conversation session according to that schedule. The conversation engine uses natural language processing to converse with the user in real time.

[0515] Examples:

[0516] At 9 a.m., the device speaks to the user, saying, "It's English time. Good morning! How did you sleep last night?" This dialogue is facilitated by generative AI, which automatically generates the next question or comment based on the user's response.

[0517] 4. Four English Skills Approach

[0518] The device offers a variety of activities to balance and develop users' listening, speaking, reading and writing skills.

[0519] Examples:

[0520] During the listening section, the device will play a short story to the user and then ask questions about the story, while during the speaking section, the device will have the user pronounce a simple phrase and evaluate their pronunciation to guide them to the next step.

[0521] 5. Mother Tongue Acquisition Program

[0522] The server periodically evaluates the user's learning progress and analyzes the progress data stored in the database, and based on this evaluation, modifies or adjusts the learning plan.

[0523] Examples:

[0524] At the end of the month, during progress assessment, the server determines that the user's listening skills have improved, but their speaking skills are lacking. Based on this, speaking-focused activities are added to the next month's learning plan.

[0525] The system of the present invention naturally integrates multilingual exposure into the user's living environment, allowing the user to naturally acquire a second language in their daily life and grow as a bilingual. This system is particularly effective in language acquisition from an early age, and solves the problems facing the modern education system.

[0526] The processing flow will be explained below.

[0527] Processing flow

[0528] Entering user data and registering

[0529] Step 1:

[0530] Users access the application and enter information such as their name, age, language level, and learning goals on the new registration screen.

[0531] Step 2:

[0532] The terminal receives the entered user information and sends it to the server.

[0533] Step 3:

[0534] The server registers the received user information in a database and generates a user ID.

[0535] First-time setup and customization

[0536] Step 1:

[0537] The server retrieves user information from a database and generates an initial learning plan based on this.

[0538] Step 2:

[0539] The device presents the generated learning plan to the user and asks for confirmation of the learning plan.

[0540] Step 3:

[0541] The user reviews the proposed study plan and requests adjustments if necessary.

[0542] Multilingual conversation function

[0543] Step 1:

[0544] The device starts a conversation session on a specified schedule based on the learning plan.

[0545] Step 2:

[0546] MimiTomo (robot) uses a natural language processing engine to converse with users in the selected language.

[0547] Step 3:

[0548] Users have the opportunity to converse with the robot and use language naturally.

[0549] Four English Skills Approach

[0550] Step 1:

[0551] The device generates and provides activities for practicing listening, speaking, reading and writing skills.

[0552] Step 2:

[0553] The user performs each skill practice activity provided by the terminal.

[0554] Step 3:

[0555] The terminal collects the user's activity results and sends them to the server.

[0556] Mother Tongue Acquisition Program

[0557] Step 1:

[0558] The server periodically evaluates the user's learning progress and updates the database based on the evaluation results.

[0559] Step 2:

[0560] The server analyzes the progress data and adjusts the learning plan as needed.

[0561] Step 3:

[0562] The device will present the user with a new, adjusted learning plan and begin the next learning cycle.

[0563] In this way, the system generates an individualized learning plan based on the user's information, engages in daily multilingual conversations based on that plan, and provides activities to train the listening, speaking, reading, and writing skills in a balanced manner.By continuously evaluating the user's progress, the system supports effective bilingual development.

[0564] Example 1

[0565] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0566] In multilingual learning, there is a need for systems that can take into account individual differences, provide users with a learning plan that is tailored to their needs, and adjust it as needed based on their progress. In particular, there is a lack of effective methods for balanced development of listening, speaking, reading, and writing skills. Another issue is the difficulty of effectively incorporating real-time multilingual conversation functions.

[0567] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0568] In this invention, the server includes means for inputting user information and generating a study plan based on the information, means for conducting multilingual conversations using a natural language processing engine based on the generated study plan, means for evaluating the user's progress and changing or adjusting the study plan based on the evaluation, and means for providing activities for practicing the skills of auditory comprehension, speaking, reading comprehension, and writing based on the generated study plan. This makes it possible to create and adjust a study plan suited to the user and develop each skill in a balanced manner.

[0569] "User" refers to an individual who uses the system to acquire multiple languages.

[0570] "Information" refers to data about the user such as name, age, language level, learning goals, etc.

[0571] "Study Plan" refers to a customized study schedule and content based on user information.

[0572] A "natural language processing engine" refers to technology that enables computers to understand and generate human language.

[0573] A "multilingual conversation" refers to a dialogue that takes place using more than one language.

[0574] "Progress" refers to the state of how far a user has progressed or improved on their learning plan.

[0575] "Evaluation" refers to the process of measuring a user's learning progress and analyzing the results.

[0576] "Auditory comprehension" refers to the skill of understanding language heard.

[0577] "Speech" refers to the skill of speaking a language.

[0578] "Reading comprehension" refers to the skill of understanding written language.

[0579] "Writing" refers to the skill of using language to write sentences.

[0580] "Activity" refers to specific exercises or tasks that a user performs to practice each skill.

[0581] The system of the present invention inputs user information, generates an individually customized learning plan based on that information, and conducts multilingual conversations using a generative AI model. It also evaluates progress and adjusts the learning plan based on that evaluation, enabling the development of auditory comprehension, speaking, reading comprehension, and writing skills in a balanced manner.

[0582] 1. Entering user data and registering

[0583] The server receives new user information and registers it in a database. User data includes name, age, language level, learning goals, etc.

[0584] Examples:

[0585] A user accesses the application and enters information such as name, age, language level, learning goals, etc. on the new registration screen. This information is received by the server and registered in the database.

[0586] 2. Initial setup and customization

[0587] The device generates an individual learning plan based on the user's information, which is customized according to the user's age, language level, and learning goals.

[0588] Examples:

[0589] If a 5-year-old user is enrolled, the device will generate a learning plan for young children, including games, songs, and activities using basic English words and phrases.

[0590] 3. Multilingual conversation function

[0591] The device sets a schedule based on the study plan and starts a multilingual conversation session according to that schedule. The conversation engine uses natural language processing to converse with the user in real time.

[0592] Examples:

[0593] At 9 a.m., the device speaks to the user, "It's time for English. Good morning! How did you sleep last night?" This dialogue is driven by a generative AI model, which automatically generates the next question or comment based on the user's response.

[0594] 4. Four English Skills Approach

[0595] The device offers a variety of activities to balance the user's auditory comprehension, speaking, reading, and writing skills.

[0596] Examples:

[0597] During the listening section, the device will play a short story to the user and then ask questions about the story. During the speaking section, the device will have the user pronounce simple phrases and evaluate their pronunciation to guide them to the next step.

[0598] 5. Mother Tongue Acquisition Program

[0599] The server periodically evaluates the user's learning progress and analyzes the progress data stored in the database, and based on this evaluation, modifies or adjusts the learning plan.

[0600] Examples:

[0601] At the end of the month, the server assesses the progress and finds that the user's listening skills have improved, but their speaking skills are lacking. Based on this, activities to strengthen speaking are added to the next month's lesson plan.

[0602] Through this entire system, multilingual exposure can be naturally incorporated into the user's living environment, allowing them to naturally acquire a second language in their daily lives and grow as bilinguals. This is particularly effective for language acquisition from an early age, and solves the problems facing the modern education system.

[0603] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0604] Step 1:

[0605] Entering user data and registering

[0606] Input: The user enters information such as name, age, language level, and learning goals on the application's sign-up screen.

[0607] Specific behavior: The data entered by the user is sent to the server through the input form.

[0608] Data processing: The server formats the received user data into an appropriate format (e.g., converts it to JSON format).

[0609] Output: Correctly formatted user data.

[0610] Step 2:

[0611] Receiving information and registering in the database

[0612] Input: The server receives the formatted user data.

[0613] What happens: The server checks the format of the data and issues the appropriate SQL query to store it in the database.

[0614] Data manipulation: Insert user data into the database using SQL queries.

[0615] Output: User data is registered in the database.

[0616] Step 3:

[0617] User information acquisition

[0618] Input: User data registered on the server.

[0619] Specific operation: The device obtains user data from the server through an API call.

[0620] Data processing: Converting received data into internal data structures.

[0621] Output: User data stored in the device.

[0622] Step 4:

[0623] Generate a lesson plan

[0624] Input: User data stored on the device.

[0625] What it does: The lesson plan generator generates an appropriate lesson plan based on the user's age, language level, and learning goals.

[0626] Data processing: The study plan generator analyzes user data and creates a customized study plan that is optimal for them.

[0627] Output: A customized learning plan for the user.

[0628] Step 5:

[0629] Scheduling

[0630] Input: Schedule information based on your study plan.

[0631] Specific behavior: The device uses the calendar API to automatically reserve a session for the specified date and time.

[0632] Data processing: Convert schedule information into calendar format and register it on the user's device.

[0633] Output: The configured schedule.

[0634] Step 6:

[0635] Starting a multilingual conversation session

[0636] Input: A prompt statement based on the lesson plan.

[0637] Specific operation: The device sends a prompt to the generative AI model and begins a dialogue with the user in real time.

[0638] Data processing: Input the prompt sentence into the natural language processing engine and receive a response from the generative AI model.

[0639] Output: An ongoing multilingual conversation with the user.

[0640] Step 7:

[0641] Skills Assessment

[0642] Input: User behavior data during a multilingual conversation session.

[0643] How it works: The device uses speech recognition technology and natural language processing to evaluate the user's pronunciation and responses.

[0644] Data processing: Analyze the collected data and generate metrics to assess the user's skill level.

[0645] Output: User skill assessment results.

[0646] Step 8:

[0647] Collecting and evaluating progress data

[0648] Input: Skill assessment results after each session.

[0649] Specific operation: The server collects and analyzes the user's learning progress data sent from the device.

[0650] Data processing: The collected data is input into a statistical model to evaluate progress.

[0651] Output: Progress assessment results and reports.

[0652] Step 9:

[0653] Adjusting your study plan

[0654] Input: Progress assessment results.

[0655] Specific operation: The server generates a new learning plan based on the progress assessment and distributes it to the device.

[0656] Data processing: Applying algorithms to progress data to create new, customized learning plans.

[0657] Output: A tailored learning plan.

[0658] (Application example 1)

[0659] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0660] Increasing the efficiency of multilingual learning is a challenge facing the modern education system. Furthermore, with limited options for effective language learning in busy daily lives, there is a need for learning methods that allow for effective use of time, especially while traveling. There is also a need for learning plans tailored to the age and language level of students, from children to adults, that foster a balanced development of listening, speaking, reading, and writing skills. To address these challenges, a system is needed that can provide customized learning plans and progress management tailored to each user's individual needs.

[0661] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0662] The invention includes means for inputting user information and generating a learning plan based on the information, means for conducting multilingual conversations using a natural language processing engine based on the generated learning plan, means for evaluating the user's progress and changing or adjusting the learning plan based on the evaluation, means for providing activities to practice listening, speaking, reading, and writing skills based on the generated learning plan, and means for providing multilingual learning using a display, microphone, and speaker in the vehicle.

[0663] This allows passengers to effectively utilize their time while traveling in a vehicle to study multiple languages, and enables them to efficiently develop their language skills in a balanced manner through a learning plan customized for each user.In addition, real-time dialogue and progress management can maximize the effectiveness of their learning.

[0664] "User Information" refers to personal data necessary to generate a learning plan, such as the learner's name, age, language level, and learning goals.

[0665] "Study Plan" refers to specific learning content and schedules that are individually customized based on the user's information and designed to develop listening, speaking, reading, and writing skills in a balanced manner.

[0666] A "natural language processing engine" refers to a software system that uses generative AI to conduct natural conversations in multiple languages.

[0667] "Progress assessment" refers to the process of quantitatively and qualitatively evaluating the results of the learning activities that a user has carried out based on their learning plan.

[0668] "Listening, speaking, reading and writing skills" refers to the four basic skill areas in language learning, with corresponding practice activities.

[0669] "In-vehicle display" refers to the screen installed inside an autonomous vehicle that displays learning content and provides interactive instructions.

[0670] A "microphone" refers to an input device that picks up a user's voice and analyzes the voice in real time.

[0671] "Speaker" refers to an audio device that allows interaction between the user and the system through audio output.

[0672] A "robot device" refers to a hardware device used to interact with a user in real time, and has functions such as voice recognition and voice synthesis.

[0673] This invention is a system for providing multilingual learning within an autonomous vehicle, which has the function of generating a learning plan based on user information and conducting multilingual conversations in real time based on that plan. The system includes the following elements:

[0674] 1. Entering user data and registering

[0675] The server receives information such as the user's name, age, language level, and learning goals provided by the user via a display, smartphone, or other device when they board an autonomous vehicle, and registers this information in a database, creating a foundation for generating learning plans tailored to the user's individual needs.

[0676] 2. Create and customize your study plan

[0677] The server uses a generative AI model based on registered user information to generate individual learning plans. For example, it can provide a business English course for adults commuting to work, and a playful basic English learning plan for children. The learning plans include a balanced mix of listening, speaking, reading, and writing skills.

[0678] 3. Multilingual conversation function

[0679] The device sets a schedule based on the learning plan and starts multilingual conversation sessions according to that schedule. Using the vehicle's microphone and speaker, the device can communicate with the user in real time using a natural language processing engine, allowing users to efficiently improve their language skills while on the move.

[0680] 4. Four Skills Approach

[0681] The device offers multifaceted activities to practice listening, speaking, reading, and writing skills based on a learning plan. For example, during listening, the device plays a short story followed by questions to check comprehension. During speaking, the device evaluates the user's pronunciation in real time and provides feedback.

[0682] 5. Progress assessment and plan adjustment

[0683] The server periodically evaluates the user's progress and analyzes the progress data stored in the database. Based on this evaluation, the server changes or adjusts the learning plan to help the user improve their skills. For example, if listening skills are improving but speaking skills are lacking, the server will add activities to strengthen speaking in the next month.

[0684] Specific examples

[0685] A user boards an autonomous vehicle and registers by entering their name, age, language level, and learning goals on the display terminal. The server then generates a customized learning plan based on the user's information and initiates a multilingual conversation session through the vehicle's speakers and display.

[0686] Prompt Sentence Examples

[0687] "Enter the user's name, age, language level, and learning goals, and generate a program to charge them in the car's interface."

[0688] This system allows users to study multiple languages ​​according to an individually customized learning plan, making effective use of their time inside an autonomous vehicle. It is expected that real-time dialogue and progress management will maximize learning effectiveness.

[0689] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0690] Step 1:

[0691] When a user gets into an autonomous vehicle, they input information via a display or smartphone. This information includes their name, age, language level, and learning goals, and the input is sent to a server. The server receives this information and registers it in a database. The input of this step is user information, and the output is the user information registered in the database.

[0692] Step 2:

[0693] The server uses a generative AI model to generate individual study plans based on registered user information. The generated study plans are customized according to the user's age, language level, and learning goals. For example, an adult user aiming for business English would be provided with a study plan focused on business situations. The input for this step is user information registered in the database, and the output is a customized study plan.

[0694] Step 3:

[0695] The device sets a schedule based on the learning plan sent from the server. Specifically, it uses the vehicle's display and speaker to indicate the required activities to the user and start the learning session. For example, it sets an English listening session to start at 9:00 AM. The input of this step is the customized learning plan, and the output is the set schedule.

[0696] Step 4:

[0697] The device initiates a multilingual conversation session according to a set schedule. It collects user utterances through a microphone in the vehicle and analyzes them through a natural language processing engine. It uses a generative AI model on the analyzed data to generate an appropriate reply and respond to the user through the speaker. The input for this step is the user's utterance data, and the output is the generated response.

[0698] Step 5:

[0699] The device provides activities for practicing listening, speaking, reading, and writing skills. For example, a listening session plays a short story followed by questions to check comprehension. A speaking session allows the user to practice pronunciation and evaluate its accuracy with a speech recognition system. The input for this step is the user's activity data based on the learning plan, and the output is the evaluation result.

[0700] Step 6:

[0701] The server periodically evaluates the user's progress and analyzes the progress data stored in the database. Based on this evaluation, it modifies or adjusts the user's learning plan. For example, if the user's listening skills are improving but their speaking skills are lacking, it will add activities to strengthen their speaking skills next month. The input of this step is the progress data stored in the database, and the output is an adjusted learning plan.

[0702] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0703] The system generates a learning plan based on user information, and utilizes a natural language processing engine and an emotion engine to evaluate the user's progress and emotions while conducting multilingual conversations, dynamically adjusting the learning plan. By nurturing listening, speaking, reading, and writing skills in a balanced manner and taking into account the user's emotional state, the system provides an optimal learning experience.

[0704] 1. Entering user data and registering

[0705] The server receives new user information and registers it in a database. User data includes name, age, language level, learning goals, etc.

[0706] Examples:

[0707] A user accesses the application and enters information such as name, age, language level, learning goals, etc. on the new registration screen. This information is received by the server and registered in the database.

[0708] 2. Initial setup and customization

[0709] The device generates an individual learning plan based on the user's information, which is customized according to the user's age, language level, and learning goals. Furthermore, an emotion engine is used to configure the plan to take into account the user's emotional state.

[0710] Examples:

[0711] When a five-year-old user registers, the device generates a learning plan for young children, including games and songs, activities using basic English words and phrases, and selects content that will interest the user and analyzes it with an emotion engine.

[0712] 3. Multilingual conversation function

[0713] The device sets a schedule based on the learning plan and starts a multilingual conversation session according to that schedule. The conversation engine uses natural language processing to converse with the user in real time. Furthermore, the emotion engine analyzes the user's emotions and adjusts the content of the conversation accordingly.

[0714] Examples:

[0715] At 9 a.m., the device speaks to the user, saying, "It's time for English. Good morning! How did you sleep last night?" This dialogue is conducted by a generative AI, which automatically generates the next question or comment based on the user's response. If the user shows no interest, the emotion engine analyzes the user's facial expressions and tone of voice and changes the topic.

[0716] 4. Four English Skills Approach

[0717] The device generates and provides activities to balance listening, speaking, reading, and writing skills, and uses an emotion engine to analyze the user's emotional state and provide appropriate feedback.

[0718] Examples:

[0719] During the listening section, the device tells the user a short story and then asks them questions about it. During the speaking section, the device asks the user to pronounce simple phrases, and the device evaluates their pronunciation and guides them to the next step. An emotion engine analyzes whether the user is enjoying themselves or struggling and adjusts its approach accordingly.

[0720] 5. Mother Tongue Acquisition Program

[0721] The server periodically evaluates the user's learning progress and analyzes the progress data stored in the database. Based on this evaluation, the server changes or adjusts the learning plan. It also uses an emotion engine to consider the user's emotional state and provide an optimal plan that maintains the user's motivation to learn.

[0722] Examples:

[0723] At the end of the month, the server will assess the user's progress and determine that their listening skills have improved, but their speaking skills are lacking. Based on this, activities to strengthen speaking will be added to the next month's learning plan. Furthermore, the emotion engine will analyze the user's motivation and optimize the learning plan to increase activities that they find challenging.

[0724] This system allows users to naturally acquire a second language in their daily lives and grow as bilinguals. By combining it with an emotion engine, it is possible to grasp the user's emotional state in real time and provide an optimal learning environment. As a result, an environment is created where multilingual conversations are constantly taking place, allowing users to grow naturally in a bilingual environment.

[0725] The processing flow will be explained below.

[0726] Processing flow

[0727] Entering user data and registering

[0728] Step 1:

[0729] Users access the application and enter information such as their name, age, language level, and learning goals on the new registration screen.

[0730] Step 2:

[0731] The terminal receives the entered user information and sends it to the server.

[0732] Step 3:

[0733] The server registers the received user information in a database and generates a user ID.

[0734] First-time setup and customization

[0735] Step 1:

[0736] The server retrieves user information from a database and generates an initial learning plan based on this.

[0737] Step 2:

[0738] The device presents the generated learning plan to the user and asks for confirmation of the learning plan.

[0739] Step 3:

[0740] The user reviews the proposed study plan and requests adjustments if necessary.

[0741] Step 4:

[0742] The server adjusts the learning plan based on user feedback and determines the final plan.

[0743] Multilingual conversation function

[0744] Step 1:

[0745] The device starts a conversation session on a specified schedule based on the learning plan.

[0746] Step 2:

[0747] MimiTomo (robot) uses a natural language processing engine to converse with users in the selected language.

[0748] Step 3:

[0749] The emotion engine analyzes the user's facial expressions and tone of voice to determine the user's emotional state.

[0750] Step 4:

[0751] MimiTomo (the robot) adjusts the conversation based on the user's emotional state, for example, changing the topic if the user is losing interest.

[0752] Step 5:

[0753] Users have the opportunity to converse with the robot and use language naturally.

[0754] Four English Skills Approach

[0755] Step 1:

[0756] The device generates and provides activities for practicing listening, speaking, reading and writing skills.

[0757] Step 2:

[0758] The user performs each skill practice activity provided by the terminal.

[0759] Step 3:

[0760] The emotion engine analyzes the user's emotional state during an activity and adjusts the difficulty and content of the activity if motivation is declining.

[0761] Step 4:

[0762] The terminal collects the user's activity results and sends them to the server.

[0763] Mother Tongue Acquisition Program

[0764] Step 1:

[0765] The server periodically evaluates the user's learning progress and updates the database based on the evaluation results.

[0766] Step 2:

[0767] The server analyzes the progress data and adjusts the learning plan as needed.

[0768] Step 3:

[0769] The emotion engine also evaluates the user's emotional data and reflects this in adjusting the learning plan.

[0770] Step 4:

[0771] The device will present the user with a new, adjusted learning plan and begin the next learning cycle.

[0772] Specific examples

[0773] Entering user data and registering

[0774] Step 1:

[0775] A user accesses the application and registers by entering their name "Taro," age "5 years old," language level "Beginner," and learning goal "Bilingual."

[0776] Step 2:

[0777] The terminal receives the input information and transmits it to the server.

[0778] Step 3:

[0779] The server registers the received information in the database and generates a user ID "001."

[0780] First-time setup and customization

[0781] Step 1:

[0782] The server generates an individual learning plan, "English Learning Plan for Toddlers," based on the user ID "001."

[0783] Step 2:

[0784] The device will present the generated learning plan to the user and ask them to "confirm."

[0785] Step 3:

[0786] The user reviews the plan and requests any necessary adjustments (e.g., adding specific activities).

[0787] Step 4:

[0788] The server adjusts the learning plan based on the request and finalizes the plan.

[0789] Multilingual conversation function

[0790] Step 1:

[0791] The device will set up and start an English conversation session at 9am every morning.

[0792] Step 2:

[0793] MimiTomo (robot) says, "Good morning, Taro! How did you sleep last night?"

[0794] Step 3:

[0795] The emotion engine analyzes the user's facial expression and determines that they are "showing interest."

[0796] Step 4:

[0797] MimiTomo (robot) continues the topic by saying, "Let's talk about your favorite toy."

[0798] Step 5:

[0799] The user responds in English, "My favorite toy is a car."

[0800] Four English Skills Approach

[0801] Step 1:

[0802] The device generates a listening activity and plays a short story.

[0803] Step 2:

[0804] Users listen to a story and answer related questions.

[0805] Step 3:

[0806] The emotion engine analyzes the user's level of concentration and determines that they are "concentrated."

[0807] Step 4:

[0808] The device provides additional activity questions to promote deeper understanding.

[0809] Mother Tongue Acquisition Program

[0810] Step 1:

[0811] The server evaluates the user's progress at the end of each month and updates the database.

[0812] Step 2:

[0813] The server analyzes the progress data and determines that "speaking skills need to be improved."

[0814] Step 3:

[0815] The emotion engine analyzes the user's motivation data and selects activities that they find rewarding.

[0816] Step 4:

[0817] The device will present the user with a new, adjusted learning plan and begin the next learning cycle.

[0818] Example 2

[0819] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0820] Conventional educational systems have had difficulty dynamically adjusting learning plans based on each learner's progress and emotional state. Especially in multilingual conversation learning, there has been a lack of appropriate feedback and activities that take into account the learner's emotional state. As a result, it has been difficult to maintain learner motivation and acquire skills efficiently.

[0821] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0822] In this invention, the server includes means for inputting user information and generating a study plan based on the information, means for conducting multilingual conversations using a natural language processing engine based on the generated study plan, means for evaluating the user's progress and changing or adjusting the study plan based on the evaluation, means for training the listening, speaking, reading, and writing skills in a balanced manner, and means for analyzing the user's emotional state and dynamically adjusting the study plan, thereby making it possible to provide an optimal learning experience according to the progress and emotional state of each individual learner.

[0823] "Means for inputting user information" refers to a method or device by which a user inputs information such as name, age, language level, learning goals, etc. into the system.

[0824] A "means for generating a study plan" is a method or device for creating an individualized study plan based on user information.

[0825] A "natural language processing engine" is a technology that uses artificial intelligence to analyze user input and generate appropriate responses.

[0826] A "means for conducting a multilingual conversation" is a method or device for conducting a conversation with a user in multiple languages.

[0827] A "means for assessing user progress" is a method or device for measuring and analyzing a user's learning progress.

[0828] A "means for changing or adjusting a study plan" is a method or device for modifying an existing study plan based on a user's progress or needs.

[0829] "A means for balanced training of listening, speaking, reading, and writing skills" is a method or device for developing these language skills in a balanced manner.

[0830] The "means for analyzing the user's emotional state" refers to a method or device for analyzing the user's facial expression, tone of voice, etc. to determine their emotions.

[0831] A "means for dynamically adjusting a study plan" is a method or apparatus for optimizing a study plan in real time according to a user's emotional state and progress.

[0832] An "information processing device" is a computer system that performs multilingual conversations, assessment of learning progress, analysis of emotional states, and so on.

[0833] This invention relates to a system that generates a study plan based on user information, and dynamically adjusts the study plan by evaluating the user's progress and emotions while conducting multilingual conversations using a natural language processing engine and an emotion engine. This system is configured to effectively support the user's language learning.

[0834] The system mainly consists of a server and a terminal. The server receives user information and registers it in a database. User data includes name, age, language level, learning goals, etc. The server accesses the database and generates an individualized initial learning plan based on the user information. The terminal receives this learning plan and configures it to predict the user's emotional state. At this point, an emotional engine (e.g., Affectiva) is used to analyze the user's emotional state.

[0835] The learning plan is designed to balance the training of listening, speaking, reading, and writing skills. The device will speak to the user according to a set schedule, for example at 9:00 AM, asking, "Good morning! How did you sleep last night?" This dialogue is conducted using a natural language processing engine (e.g., GPT-3), which generates an appropriate response based on the user's input. If the user shows no interest, the emotion engine analyzes the user's facial expressions and tone of voice and changes the topic.

[0836] The device also offers listening, speaking, reading, and writing activities. For example, during listening, it plays a short story followed by questions about the story. When the user answers the questions, the device evaluates the answer and guides the user to the next step. The emotion engine analyzes the user's reaction and provides appropriate feedback.

[0837] The server periodically retrieves the user's learning progress data from the database and evaluates that progress. For example, if listening skills are improving but speaking skills are lacking, activities to strengthen speaking will be added to the next month's learning plan. The emotion engine also analyzes the user's motivation and optimizes the plan to increase activities that are likely to be rewarding. The server then sends the updated learning plan to the device, allowing for a smooth start to the next month's learning.

[0838] This system allows users to learn multiple languages ​​naturally in their daily lives, and by combining it with an emotion engine, it is possible to grasp the user's emotional state in real time and provide the optimal learning environment. Multilingual conversations utilizing generative AI models and prompt sentences allow users to constantly have new learning experiences. By utilizing technologies such as "generative AI models and prompt sentences," the system provides the most effective learning method for users.

[0839] Examples:

[0840] Example prompt: "Good morning! How did you sleep last night?" The next question or comment is automatically generated based on the user's response.

[0841] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0842] System program processing steps

[0843] Step 1: Entering user data and registering

[0844] 1-1: The user opens the new registration screen.

[0845] Input: The user enters information such as name, age, language level, and learning goals.

[0846] 1-2: The server receives the information sent by the user.

[0847] Data processing: Convert the received user information into database format.

[0848] 1-3: The server registers the information in the database.

[0849] Output: Generates a message confirming successful registration and sends it to the terminal.

[0850] What it does: The server uses a database management system (e.g., MySQL) to store data.

[0851] Step 2: Initial setup and customization

[0852] 2-1: The server reads the user information in the database.

[0853] Input: User information retrieved from the database.

[0854] 2-2: The server analyzes the user information and generates an initial learning plan.

[0855] Data processing: Generate a customized learning plan based on the user's age, language level, and learning goals.

[0856] 2-3: The device is configured to predict the user's emotional state based on the learning plan received from the server.

[0857] Input: The learning plan sent by the server.

[0858] 2-4: The device presents the learning plan to the user and asks for confirmation.

[0859] Output: Generates a preview of the learning plan and displays it to the user.

[0860] Specific operation: The device configures the emotion engine (e.g., Affectiva) and prepares it to analyze the user's emotional state.

[0861] Step 3: Multilingual conversation function

[0862] 3-1: The device starts a conversation session according to the set schedule.

[0863] Input: Schedule information based on your study plan.

[0864] 3-2: The user talks to the terminal.

[0865] Input: User voice or text input.

[0866] 3-3: The terminal analyzes the user's input using a natural language processing engine.

[0867] Data Computing: Using a natural language processing engine (e.g., GPT-3) to analyze user input and generate appropriate responses.

[0868] 3-4: The device uses an emotion engine to analyze the user's emotional state and dynamically adjust the content of the conversation.

[0869] Input: User's facial expressions and tone of voice.

[0870] Output: Generates tailored conversation content and presents it to the user.

[0871] Specific operation: The device responds to the user with the conversation content generated by the device via voice or text.

[0872] Step 4: Four English Skills Approach

[0873] 4-1: The device starts the learning activity at the set time.

[0874] Input: Activity schedule based on the lesson plan.

[0875] 4-2: The devices provide listening, speaking, reading, and writing activities.

[0876] Input: User actions or responses.

[0877] Data processing: Record and analyze user responses to each activity.

[0878] 4-3: The device uses an emotion engine to evaluate the user's reaction and provide feedback as needed.

[0879] Input: The user's emotional state.

[0880] Output: Generate feedback based on the evaluation results.

[0881] What it does: The listening activity plays a short story, and the speaking activity teaches the user pronunciation.

[0882] Step 5: Mother Tongue Acquisition Program

[0883] 5-1: The server periodically retrieves the user's learning progress data from the database.

[0884] Input: Progress data accumulated in the database.

[0885] 5-2: The server analyzes the progress data and updates the learning plan.

[0886] Data calculations: Analyze progress data to identify skills that users need to improve.

[0887] 5-3: The server uses the emotion engine to evaluate the user's emotional state and optimize the learning plan.

[0888] Input: User sentiment analysis data.

[0889] Output: Generates an optimized learning plan and sends it to the device.

[0890] 5-4: The server sends the updated learning plan to the device.

[0891] Output: Sends the updated study plan to the device, ready to start studying for the next month.

[0892] What it does: The server uses analytics tools to evaluate progress data and generate a learning plan with appropriate changes.

[0893] (Application example 2)

[0894] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0895] Conventional multilingual conversation systems lack the ability to dynamically analyze and respond to users' progress and emotional state in real time, resulting in insufficient learning efficiency and user interaction. Furthermore, they lacked the information processing capabilities to provide personalized responses when dealing with customers in physical stores. This can result in a lack of motivation to learn and a decline in user satisfaction. In particular, flexible responses that take into account the user's emotional state are essential for complex multilingual interactions.

[0896] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0897] In this invention, the server includes: means for inputting user information and generating a study plan based on the information; means for conducting multilingual conversations using a natural language processing engine based on the generated study plan; means for evaluating the user's progress and changing or adjusting the study plan based on the evaluation; means for providing activities for practicing listening, speaking, reading, and writing skills based on the generated study plan; means including an emotion analysis engine for analyzing the user's emotional state and providing an optimal response in real time; and means including an information processing device for responding to customers in a store. This makes it possible to analyze the user's study progress and emotional state in real time and provide an optimal study plan based on the analysis, while also enabling personalized and flexible customer service in physical stores.

[0898] "User" refers to an individual who uses the system.

[0899] "Information" refers to data about the user such as name, age, language level, learning goals, etc.

[0900] "Study Plan" refers to a schedule and content of study activities customized based on user information.

[0901] A "natural language processing engine" refers to a software component that analyzes text data and conducts conversations in multiple languages.

[0902] "Multilingual conversation" refers to interaction using multiple languages, during which real-time translation and dialogue control are carried out.

[0903] "Progress" refers to the results and achievement of a user's learning activities.

[0904] "Listening" refers to the skill of understanding spoken information.

[0905] "Speaking" refers to the skill of conveying information orally.

[0906] "Reading" refers to the skill of reading and understanding written information.

[0907] "Writing" refers to the skill of describing information using written text.

[0908] "Emotional state" refers to a user's current mental and emotional state.

[0909] An "emotion analysis engine" refers to a software component that analyzes a user's emotional state and responds appropriately.

[0910] "Information processing device" refers to a device that includes hardware and software for inputting, processing, and outputting data.

[0911] "Store" means a physical location for conducting sales activities.

[0912] This system starts by inputting user information and generating a learning plan based on that information. Specifically, the server receives information such as the user's name, age, language level, and learning goals, and registers it in a pre-prepared database. This forms the basis for customizing an appropriate learning plan for each user.

[0913] Next, the device uses a natural language processing engine to conduct a multilingual conversation based on the generated learning plan. Specifically, a schedule is created based on the information entered by the user, and the conversation proceeds according to that schedule. During the conversation, an emotion analysis engine analyzes the user's emotional state in real time, and the content of the conversation is adjusted accordingly.

[0914] The system constantly evaluates the user's progress and changes or adjusts the learning plan based on the evaluation results. This method makes it possible to provide an optimal learning plan that is always up to date. For example, if the user's listening skills improve, new skill practice will be incorporated to advance to the next level.

[0915] Furthermore, the system provides activities to practice listening, speaking, reading, and writing skills in a balanced manner. During the activities, the system analyzes the user's emotional state and provides appropriate feedback. This analysis is performed using software components such as EmotionAnalyzer and NLPModel.

[0916] A specific example is an information processing device installed in a cash register that interacts with the user in real time. For example, a smartphone or information robot can provide personalized guidance based on the user's name and learning goals. If the user appears unhappy, an emotion analysis engine can analyze their emotional state and flexibly adapt the response to improve user satisfaction.

[0917] An example of a prompt is as follows:

[0918] "Customer ID: 1

[0919] Customer Name: Sato

[0920] Language: Japanese

[0921] Age: 30

[0922] Emotional state: happy

[0923] Response: Personalized greetings and product recommendations

[0924] Output: 'Hello, Sato-san. How can I help you today? We have great deals on electronics today.'

[0925] Based on this prompt, the generative AI model generates appropriate answers and responses, making customer service in physical stores much smoother.

[0926] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0927] Step 1:

[0928] The server inputs user information and registers it in a database. The input information includes name, age, language level, learning goals, etc. Specifically, when a user inputs information using a smartphone or interface device, the server receives it and stores it in a database.

[0929] Input: User's name, age, language level, learning goal

[0930] Output: User information registered in the database

[0931] Step 2:

[0932] The server generates a learning plan based on the information stored in the database, including content tailored to the user's age, language level, and learning goals, and uses a sentiment analysis engine to customize the plan based on the user's emotional state.

[0933] Input: User information registered in the database

[0934] Output: A customized study plan

[0935] Step 3:

[0936] The device initiates a multilingual conversation using a natural language processing engine based on the generated learning plan. Specifically, the device proceeds with a dialogue with the user according to the learning plan and generates appropriate responses in real time using a generative AI model.

[0937] Input: Customized Study Plan

[0938] Output: User interaction, real-time generated responses

[0939] Step 4:

[0940] The server periodically evaluates the user's progress and modifies or adjusts the learning plan based on the evaluation, which includes activity log data and user feedback.

[0941] Input: Learning activity log data, user feedback

[0942] Output: Adjusted learning plan

[0943] Step 5:

[0944] The device performs activities to practice listening, speaking, reading and writing skills based on a tailored learning plan, and during each activity, an emotion analysis engine analyzes the user's emotional state and provides appropriate feedback.

[0945] Input: Adjusted Study Plan

[0946] Output: User's emotional state data, feedback on activities

[0947] Step 6:

[0948] The server responds to users using the in-store information processing device. Specifically, when a user makes an inquiry to the in-store information processing device, the emotion analysis engine analyzes the user's emotional state based on the information and generates an appropriate response. Based on example prompt sentences, the generative AI model provides an appropriate response to the inquiry.

[0949] Input: In-store inquiry details, user's emotional state

[0950] Output: Personalized query response

[0951] Through the above processing steps, the system analyzes the user's learning progress and emotional state in real time, providing an optimal learning plan, and enabling personalized and flexible customer service in physical stores.

[0952] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0953] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0954] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0955] [Third embodiment]

[0956] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0957] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0958] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0959] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0960] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0961] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0962] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0963] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0964] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0965] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0966] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0967] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0968] The system of the present invention allows users to input their information, generate an individually customized learning plan based on that information, and engage in multilingual conversations using a generation AI. It also evaluates progress and adjusts the learning plan based on that evaluation, enabling the development of listening, speaking, reading, and writing skills in a balanced manner.

[0969] 1. Entering user data and registering

[0970] The server receives new user information and registers it in a database. User data includes name, age, language level, learning goals, etc.

[0971] Examples:

[0972] A user accesses the application and enters information such as name, age, language level, learning goals, etc. on the new registration screen. This information is received by the server and registered in the database.

[0973] 2. Initial setup and customization

[0974] The device generates an individual learning plan based on the user's information, which is customized according to the user's age, language level, and learning goals.

[0975] Examples:

[0976] When a 5-year-old user is enrolled, the device generates a learning plan for young children, including games, songs, and activities using basic English words and phrases.

[0977] 3. Multilingual conversation function

[0978] The device sets a schedule based on the learning plan and starts a multilingual conversation session according to that schedule. The conversation engine uses natural language processing to converse with the user in real time.

[0979] Examples:

[0980] At 9 a.m., the device speaks to the user, saying, "It's English time. Good morning! How did you sleep last night?" This dialogue is facilitated by generative AI, which automatically generates the next question or comment based on the user's response.

[0981] 4. Four English Skills Approach

[0982] The device offers a variety of activities to balance and develop users' listening, speaking, reading and writing skills.

[0983] Examples:

[0984] During the listening section, the device will play a short story to the user and then ask questions about the story, while during the speaking section, the device will have the user pronounce a simple phrase and evaluate their pronunciation to guide them to the next step.

[0985] 5. Mother Tongue Acquisition Program

[0986] The server periodically evaluates the user's learning progress and analyzes the progress data stored in the database, and based on this evaluation, modifies or adjusts the learning plan.

[0987] Examples:

[0988] At the end of the month, during progress assessment, the server determines that the user's listening skills have improved, but their speaking skills are lacking. Based on this, speaking-focused activities are added to the next month's learning plan.

[0989] The system of the present invention naturally integrates multilingual exposure into the user's living environment, allowing the user to naturally acquire a second language in their daily life and grow as a bilingual. This system is particularly effective in language acquisition from an early age, and solves the problems facing the modern education system.

[0990] The processing flow will be explained below.

[0991] Processing flow

[0992] Entering user data and registering

[0993] Step 1:

[0994] Users access the application and enter information such as their name, age, language level, and learning goals on the new registration screen.

[0995] Step 2:

[0996] The terminal receives the entered user information and sends it to the server.

[0997] Step 3:

[0998] The server registers the received user information in a database and generates a user ID.

[0999] First-time setup and customization

[1000] Step 1:

[1001] The server retrieves user information from a database and generates an initial learning plan based on this.

[1002] Step 2:

[1003] The device presents the generated learning plan to the user and asks for confirmation of the learning plan.

[1004] Step 3:

[1005] The user reviews the proposed study plan and requests adjustments if necessary.

[1006] Multilingual conversation function

[1007] Step 1:

[1008] The device starts a conversation session on a specified schedule based on the learning plan.

[1009] Step 2:

[1010] MimiTomo (robot) uses a natural language processing engine to converse with users in the selected language.

[1011] Step 3:

[1012] Users have the opportunity to converse with the robot and use language naturally.

[1013] Four English Skills Approach

[1014] Step 1:

[1015] The device generates and provides activities for practicing listening, speaking, reading and writing skills.

[1016] Step 2:

[1017] The user performs each skill practice activity provided by the terminal.

[1018] Step 3:

[1019] The terminal collects the user's activity results and sends them to the server.

[1020] Mother Tongue Acquisition Program

[1021] Step 1:

[1022] The server periodically evaluates the user's learning progress and updates the database based on the evaluation results.

[1023] Step 2:

[1024] The server analyzes the progress data and adjusts the learning plan as needed.

[1025] Step 3:

[1026] The device will present the user with a new, adjusted learning plan and begin the next learning cycle.

[1027] In this way, the system generates an individualized learning plan based on the user's information, engages in daily multilingual conversations based on that plan, and provides activities to train the listening, speaking, reading, and writing skills in a balanced manner.By continuously evaluating the user's progress, the system supports effective bilingual development.

[1028] Example 1

[1029] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1030] In multilingual learning, there is a need for systems that can take into account individual differences, provide users with a learning plan that is tailored to their needs, and adjust it as needed based on their progress. In particular, there is a lack of effective methods for balanced development of listening, speaking, reading, and writing skills. Another issue is the difficulty of effectively incorporating real-time multilingual conversation functions.

[1031] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1032] In this invention, the server includes means for inputting user information and generating a study plan based on the information, means for conducting multilingual conversations using a natural language processing engine based on the generated study plan, means for evaluating the user's progress and changing or adjusting the study plan based on the evaluation, and means for providing activities for practicing the skills of auditory comprehension, speaking, reading comprehension, and writing based on the generated study plan. This makes it possible to create and adjust a study plan suited to the user and develop each skill in a balanced manner.

[1033] "User" refers to an individual who uses the system to acquire multiple languages.

[1034] "Information" refers to data about the user such as name, age, language level, learning goals, etc.

[1035] "Study Plan" refers to a customized study schedule and content based on user information.

[1036] A "natural language processing engine" refers to technology that enables computers to understand and generate human language.

[1037] A "multilingual conversation" refers to a dialogue that takes place using more than one language.

[1038] "Progress" refers to the state of how far a user has progressed or improved on their learning plan.

[1039] "Evaluation" refers to the process of measuring a user's learning progress and analyzing the results.

[1040] "Auditory comprehension" refers to the skill of understanding language heard.

[1041] "Speech" refers to the skill of speaking a language.

[1042] "Reading comprehension" refers to the skill of understanding written language.

[1043] "Writing" refers to the skill of using language to write sentences.

[1044] "Activity" refers to specific exercises or tasks that a user performs to practice each skill.

[1045] The system of the present invention inputs user information, generates an individually customized learning plan based on that information, and conducts multilingual conversations using a generative AI model. It also evaluates progress and adjusts the learning plan based on that evaluation, enabling the development of auditory comprehension, speaking, reading comprehension, and writing skills in a balanced manner.

[1046] 1. Entering user data and registering

[1047] The server receives new user information and registers it in a database. User data includes name, age, language level, learning goals, etc.

[1048] Examples:

[1049] A user accesses the application and enters information such as name, age, language level, learning goals, etc. on the new registration screen. This information is received by the server and registered in the database.

[1050] 2. Initial setup and customization

[1051] The device generates an individual learning plan based on the user's information, which is customized according to the user's age, language level, and learning goals.

[1052] Examples:

[1053] If a 5-year-old user is enrolled, the device will generate a learning plan for young children, including games, songs, and activities using basic English words and phrases.

[1054] 3. Multilingual conversation function

[1055] The device sets a schedule based on the study plan and starts a multilingual conversation session according to that schedule. The conversation engine uses natural language processing to converse with the user in real time.

[1056] Examples:

[1057] At 9 a.m., the device speaks to the user, "It's time for English. Good morning! How did you sleep last night?" This dialogue is driven by a generative AI model, which automatically generates the next question or comment based on the user's response.

[1058] 4. Four English Skills Approach

[1059] The device offers a variety of activities to balance the user's auditory comprehension, speaking, reading, and writing skills.

[1060] Examples:

[1061] During the listening section, the device will play a short story to the user and then ask questions about the story. During the speaking section, the device will have the user pronounce simple phrases and evaluate their pronunciation to guide them to the next step.

[1062] 5. Mother Tongue Acquisition Program

[1063] The server periodically evaluates the user's learning progress and analyzes the progress data stored in the database, and based on this evaluation, modifies or adjusts the learning plan.

[1064] Examples:

[1065] At the end of the month, the server assesses the progress and finds that the user's listening skills have improved, but their speaking skills are lacking. Based on this, activities to strengthen speaking are added to the next month's lesson plan.

[1066] Through this entire system, multilingual exposure can be naturally incorporated into the user's living environment, allowing them to naturally acquire a second language in their daily lives and grow as bilinguals. This is particularly effective for language acquisition from an early age, and solves the problems facing the modern education system.

[1067] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1068] Step 1:

[1069] Entering user data and registering

[1070] Input: The user enters information such as name, age, language level, and learning goals on the application's sign-up screen.

[1071] Specific behavior: The data entered by the user is sent to the server through the input form.

[1072] Data processing: The server formats the received user data into an appropriate format (e.g., converts it to JSON format).

[1073] Output: Correctly formatted user data.

[1074] Step 2:

[1075] Receiving information and registering in the database

[1076] Input: The server receives the formatted user data.

[1077] What happens: The server checks the format of the data and issues the appropriate SQL query to store it in the database.

[1078] Data manipulation: Insert user data into the database using SQL queries.

[1079] Output: User data is registered in the database.

[1080] Step 3:

[1081] User information acquisition

[1082] Input: User data registered on the server.

[1083] Specific operation: The device obtains user data from the server through an API call.

[1084] Data processing: Converting received data into internal data structures.

[1085] Output: User data stored in the device.

[1086] Step 4:

[1087] Generate a lesson plan

[1088] Input: User data stored on the device.

[1089] What it does: The lesson plan generator generates an appropriate lesson plan based on the user's age, language level, and learning goals.

[1090] Data processing: The study plan generator analyzes user data and creates a customized study plan that is optimal for them.

[1091] Output: A customized learning plan for the user.

[1092] Step 5:

[1093] Scheduling

[1094] Input: Schedule information based on your study plan.

[1095] Specific behavior: The device uses the calendar API to automatically reserve a session for the specified date and time.

[1096] Data processing: Convert schedule information into calendar format and register it on the user's device.

[1097] Output: The configured schedule.

[1098] Step 6:

[1099] Starting a multilingual conversation session

[1100] Input: A prompt statement based on the lesson plan.

[1101] Specific operation: The device sends a prompt to the generative AI model and begins a dialogue with the user in real time.

[1102] Data processing: Input the prompt sentence into the natural language processing engine and receive a response from the generative AI model.

[1103] Output: An ongoing multilingual conversation with the user.

[1104] Step 7:

[1105] Skills Assessment

[1106] Input: User behavior data during a multilingual conversation session.

[1107] How it works: The device uses speech recognition technology and natural language processing to evaluate the user's pronunciation and responses.

[1108] Data processing: Analyze the collected data and generate metrics to assess the user's skill level.

[1109] Output: User skill assessment results.

[1110] Step 8:

[1111] Collecting and evaluating progress data

[1112] Input: Skill assessment results after each session.

[1113] Specific operation: The server collects and analyzes the user's learning progress data sent from the device.

[1114] Data processing: The collected data is input into a statistical model to evaluate progress.

[1115] Output: Progress assessment results and reports.

[1116] Step 9:

[1117] Adjusting your study plan

[1118] Input: Progress assessment results.

[1119] Specific operation: The server generates a new learning plan based on the progress assessment and distributes it to the device.

[1120] Data processing: Applying algorithms to progress data to create new, customized learning plans.

[1121] Output: A tailored learning plan.

[1122] (Application example 1)

[1123] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1124] Increasing the efficiency of multilingual learning is a challenge facing the modern education system. Furthermore, with limited options for effective language learning in busy daily lives, there is a need for learning methods that allow for effective use of time, especially while traveling. There is also a need for learning plans tailored to the age and language level of students, from children to adults, that foster a balanced development of listening, speaking, reading, and writing skills. To address these challenges, a system is needed that can provide customized learning plans and progress management tailored to each user's individual needs.

[1125] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1126] The invention includes means for inputting user information and generating a learning plan based on the information, means for conducting multilingual conversations using a natural language processing engine based on the generated learning plan, means for evaluating the user's progress and changing or adjusting the learning plan based on the evaluation, means for providing activities to practice listening, speaking, reading, and writing skills based on the generated learning plan, and means for providing multilingual learning using a display, microphone, and speaker in the vehicle.

[1127] This allows passengers to effectively utilize their time while traveling in a vehicle to study multiple languages, and enables them to efficiently develop their language skills in a balanced manner through a learning plan customized for each user.In addition, real-time dialogue and progress management can maximize the effectiveness of their learning.

[1128] "User Information" refers to personal data necessary to generate a learning plan, such as the learner's name, age, language level, and learning goals.

[1129] "Study Plan" refers to specific learning content and schedules that are individually customized based on the user's information and designed to develop listening, speaking, reading, and writing skills in a balanced manner.

[1130] A "natural language processing engine" refers to a software system that uses generative AI to conduct natural conversations in multiple languages.

[1131] "Progress assessment" refers to the process of quantitatively and qualitatively evaluating the results of the learning activities that a user has carried out based on their learning plan.

[1132] "Listening, speaking, reading and writing skills" refers to the four basic skill areas in language learning, with corresponding practice activities.

[1133] "In-vehicle display" refers to the screen installed inside an autonomous vehicle that displays learning content and provides interactive instructions.

[1134] A "microphone" refers to an input device that picks up a user's voice and analyzes the voice in real time.

[1135] "Speaker" refers to an audio device that allows interaction between the user and the system through audio output.

[1136] A "robot device" refers to a hardware device used to interact with a user in real time, and has functions such as voice recognition and voice synthesis.

[1137] This invention is a system for providing multilingual learning within an autonomous vehicle, which has the function of generating a learning plan based on user information and conducting multilingual conversations in real time based on that plan. The system includes the following elements:

[1138] 1. Entering user data and registering

[1139] The server receives information such as the user's name, age, language level, and learning goals provided by the user via a display, smartphone, or other device when they board an autonomous vehicle, and registers this information in a database, creating a foundation for generating learning plans tailored to the user's individual needs.

[1140] 2. Create and customize your study plan

[1141] The server uses a generative AI model based on registered user information to generate individual learning plans. For example, it can provide a business English course for adults commuting to work, and a playful basic English learning plan for children. The learning plans include a balanced mix of listening, speaking, reading, and writing skills.

[1142] 3. Multilingual conversation function

[1143] The device sets a schedule based on the learning plan and starts multilingual conversation sessions according to that schedule. Using the vehicle's microphone and speaker, the device can communicate with the user in real time using a natural language processing engine, allowing users to efficiently improve their language skills while on the move.

[1144] 4. Four Skills Approach

[1145] The device offers multifaceted activities to practice listening, speaking, reading, and writing skills based on a learning plan. For example, during listening, the device plays a short story followed by questions to check comprehension. During speaking, the device evaluates the user's pronunciation in real time and provides feedback.

[1146] 5. Progress assessment and plan adjustment

[1147] The server periodically evaluates the user's progress and analyzes the progress data stored in the database. Based on this evaluation, the server changes or adjusts the learning plan to help the user improve their skills. For example, if listening skills are improving but speaking skills are lacking, the server will add activities to strengthen speaking in the next month.

[1148] Specific examples

[1149] A user boards an autonomous vehicle and registers by entering their name, age, language level, and learning goals on the display terminal. The server then generates a customized learning plan based on the user's information and initiates a multilingual conversation session through the vehicle's speakers and display.

[1150] Prompt Sentence Examples

[1151] "Enter the user's name, age, language level, and learning goals, and generate a program to charge them in the car's interface."

[1152] This system allows users to study multiple languages ​​according to an individually customized learning plan, making effective use of their time inside an autonomous vehicle. It is expected that real-time dialogue and progress management will maximize learning effectiveness.

[1153] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1154] Step 1:

[1155] When a user gets into an autonomous vehicle, they input information via a display or smartphone. This information includes their name, age, language level, and learning goals, and the input is sent to a server. The server receives this information and registers it in a database. The input of this step is user information, and the output is the user information registered in the database.

[1156] Step 2:

[1157] The server uses a generative AI model to generate individual study plans based on registered user information. The generated study plans are customized according to the user's age, language level, and learning goals. For example, an adult user aiming for business English would be provided with a study plan focused on business situations. The input for this step is user information registered in the database, and the output is a customized study plan.

[1158] Step 3:

[1159] The device sets a schedule based on the learning plan sent from the server. Specifically, it uses the vehicle's display and speaker to indicate the required activities to the user and start the learning session. For example, it sets an English listening session to start at 9:00 AM. The input of this step is the customized learning plan, and the output is the set schedule.

[1160] Step 4:

[1161] The device initiates a multilingual conversation session according to a set schedule. It collects user utterances through a microphone in the vehicle and analyzes them through a natural language processing engine. It uses a generative AI model on the analyzed data to generate an appropriate reply and respond to the user through the speaker. The input for this step is the user's utterance data, and the output is the generated response.

[1162] Step 5:

[1163] The device provides activities for practicing listening, speaking, reading, and writing skills. For example, a listening session plays a short story followed by questions to check comprehension. A speaking session allows the user to practice pronunciation and evaluate its accuracy with a speech recognition system. The input for this step is the user's activity data based on the learning plan, and the output is the evaluation result.

[1164] Step 6:

[1165] The server periodically evaluates the user's progress and analyzes the progress data stored in the database. Based on this evaluation, it modifies or adjusts the user's learning plan. For example, if the user's listening skills are improving but their speaking skills are lacking, it will add activities to strengthen their speaking skills next month. The input of this step is the progress data stored in the database, and the output is an adjusted learning plan.

[1166] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1167] The system generates a learning plan based on user information, and utilizes a natural language processing engine and an emotion engine to evaluate the user's progress and emotions while conducting multilingual conversations, dynamically adjusting the learning plan. By nurturing listening, speaking, reading, and writing skills in a balanced manner and taking into account the user's emotional state, the system provides an optimal learning experience.

[1168] 1. Entering user data and registering

[1169] The server receives new user information and registers it in a database. User data includes name, age, language level, learning goals, etc.

[1170] Examples:

[1171] A user accesses the application and enters information such as name, age, language level, learning goals, etc. on the new registration screen. This information is received by the server and registered in the database.

[1172] 2. Initial setup and customization

[1173] The device generates an individual learning plan based on the user's information, which is customized according to the user's age, language level, and learning goals. Furthermore, an emotion engine is used to configure the plan to take into account the user's emotional state.

[1174] Examples:

[1175] When a five-year-old user registers, the device generates a learning plan for young children, including games and songs, activities using basic English words and phrases, and selects content that will interest the user and analyzes it with an emotion engine.

[1176] 3. Multilingual conversation function

[1177] The device sets a schedule based on the learning plan and starts a multilingual conversation session according to that schedule. The conversation engine uses natural language processing to converse with the user in real time. Furthermore, the emotion engine analyzes the user's emotions and adjusts the content of the conversation accordingly.

[1178] Examples:

[1179] At 9 a.m., the device speaks to the user, saying, "It's time for English. Good morning! How did you sleep last night?" This dialogue is conducted by a generative AI, which automatically generates the next question or comment based on the user's response. If the user shows no interest, the emotion engine analyzes the user's facial expressions and tone of voice and changes the topic.

[1180] 4. Four English Skills Approach

[1181] The device generates and provides activities to balance listening, speaking, reading, and writing skills, and uses an emotion engine to analyze the user's emotional state and provide appropriate feedback.

[1182] Examples:

[1183] During the listening section, the device tells the user a short story and then asks them questions about it. During the speaking section, the device asks the user to pronounce simple phrases, and the device evaluates their pronunciation and guides them to the next step. An emotion engine analyzes whether the user is enjoying themselves or struggling and adjusts its approach accordingly.

[1184] 5. Mother Tongue Acquisition Program

[1185] The server periodically evaluates the user's learning progress and analyzes the progress data stored in the database. Based on this evaluation, the server changes or adjusts the learning plan. It also uses an emotion engine to consider the user's emotional state and provide an optimal plan that maintains the user's motivation to learn.

[1186] Examples:

[1187] At the end of the month, the server will assess the user's progress and determine that their listening skills have improved, but their speaking skills are lacking. Based on this, activities to strengthen speaking will be added to the next month's learning plan. Furthermore, the emotion engine will analyze the user's motivation and optimize the learning plan to increase activities that they find challenging.

[1188] This system allows users to naturally acquire a second language in their daily lives and grow as bilinguals. By combining it with an emotion engine, it is possible to grasp the user's emotional state in real time and provide an optimal learning environment. As a result, an environment is created where multilingual conversations are constantly taking place, allowing users to grow naturally in a bilingual environment.

[1189] The processing flow will be explained below.

[1190] Processing flow

[1191] Entering user data and registering

[1192] Step 1:

[1193] Users access the application and enter information such as their name, age, language level, and learning goals on the new registration screen.

[1194] Step 2:

[1195] The terminal receives the entered user information and sends it to the server.

[1196] Step 3:

[1197] The server registers the received user information in a database and generates a user ID.

[1198] First-time setup and customization

[1199] Step 1:

[1200] The server retrieves user information from a database and generates an initial learning plan based on this.

[1201] Step 2:

[1202] The device presents the generated learning plan to the user and asks for confirmation of the learning plan.

[1203] Step 3:

[1204] The user reviews the proposed study plan and requests adjustments if necessary.

[1205] Step 4:

[1206] The server adjusts the learning plan based on user feedback and determines the final plan.

[1207] Multilingual conversation function

[1208] Step 1:

[1209] The device starts a conversation session on a specified schedule based on the learning plan.

[1210] Step 2:

[1211] MimiTomo (robot) uses a natural language processing engine to converse with users in the selected language.

[1212] Step 3:

[1213] The emotion engine analyzes the user's facial expressions and tone of voice to determine the user's emotional state.

[1214] Step 4:

[1215] MimiTomo (the robot) adjusts the conversation based on the user's emotional state, for example, changing the topic if the user is losing interest.

[1216] Step 5:

[1217] Users have the opportunity to converse with the robot and use language naturally.

[1218] Four English Skills Approach

[1219] Step 1:

[1220] The device generates and provides activities for practicing listening, speaking, reading and writing skills.

[1221] Step 2:

[1222] The user performs each skill practice activity provided by the terminal.

[1223] Step 3:

[1224] The emotion engine analyzes the user's emotional state during an activity and adjusts the difficulty and content of the activity if motivation is declining.

[1225] Step 4:

[1226] The terminal collects the user's activity results and sends them to the server.

[1227] Mother Tongue Acquisition Program

[1228] Step 1:

[1229] The server periodically evaluates the user's learning progress and updates the database based on the evaluation results.

[1230] Step 2:

[1231] The server analyzes the progress data and adjusts the learning plan as needed.

[1232] Step 3:

[1233] The emotion engine also evaluates the user's emotional data and reflects this in adjusting the learning plan.

[1234] Step 4:

[1235] The device will present the user with a new, adjusted learning plan and begin the next learning cycle.

[1236] Specific examples

[1237] Entering user data and registering

[1238] Step 1:

[1239] A user accesses the application and registers by entering their name "Taro," age "5 years old," language level "Beginner," and learning goal "Bilingual."

[1240] Step 2:

[1241] The terminal receives the input information and transmits it to the server.

[1242] Step 3:

[1243] The server registers the received information in the database and generates a user ID "001."

[1244] First-time setup and customization

[1245] Step 1:

[1246] The server generates an individual learning plan, "English Learning Plan for Toddlers," based on the user ID "001."

[1247] Step 2:

[1248] The device will present the generated learning plan to the user and ask them to "confirm."

[1249] Step 3:

[1250] The user reviews the plan and requests any necessary adjustments (e.g., adding specific activities).

[1251] Step 4:

[1252] The server adjusts the learning plan based on the request and finalizes the plan.

[1253] Multilingual conversation function

[1254] Step 1:

[1255] The device will set up and start an English conversation session at 9am every morning.

[1256] Step 2:

[1257] MimiTomo (robot) says, "Good morning, Taro! How did you sleep last night?"

[1258] Step 3:

[1259] The emotion engine analyzes the user's facial expression and determines that they are "showing interest."

[1260] Step 4:

[1261] MimiTomo (robot) continues the topic by saying, "Let's talk about your favorite toy."

[1262] Step 5:

[1263] The user responds in English, "My favorite toy is a car."

[1264] Four English Skills Approach

[1265] Step 1:

[1266] The device generates a listening activity and plays a short story.

[1267] Step 2:

[1268] Users listen to a story and answer related questions.

[1269] Step 3:

[1270] The emotion engine analyzes the user's level of concentration and determines that they are "concentrated."

[1271] Step 4:

[1272] The device provides additional activity questions to promote deeper understanding.

[1273] Mother Tongue Acquisition Program

[1274] Step 1:

[1275] The server evaluates the user's progress at the end of each month and updates the database.

[1276] Step 2:

[1277] The server analyzes the progress data and determines that "speaking skills need to be improved."

[1278] Step 3:

[1279] The emotion engine analyzes the user's motivation data and selects activities that they find rewarding.

[1280] Step 4:

[1281] The device will present the user with a new, adjusted learning plan and begin the next learning cycle.

[1282] Example 2

[1283] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1284] Conventional educational systems have had difficulty dynamically adjusting learning plans based on each learner's progress and emotional state. Especially in multilingual conversation learning, there has been a lack of appropriate feedback and activities that take into account the learner's emotional state. As a result, it has been difficult to maintain learner motivation and acquire skills efficiently.

[1285] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1286] In this invention, the server includes means for inputting user information and generating a study plan based on the information, means for conducting multilingual conversations using a natural language processing engine based on the generated study plan, means for evaluating the user's progress and changing or adjusting the study plan based on the evaluation, means for training the listening, speaking, reading, and writing skills in a balanced manner, and means for analyzing the user's emotional state and dynamically adjusting the study plan, thereby making it possible to provide an optimal learning experience according to the progress and emotional state of each individual learner.

[1287] "Means for inputting user information" refers to a method or device by which a user inputs information such as name, age, language level, learning goals, etc. into the system.

[1288] A "means for generating a study plan" is a method or device for creating an individualized study plan based on user information.

[1289] A "natural language processing engine" is a technology that uses artificial intelligence to analyze user input and generate appropriate responses.

[1290] A "means for conducting a multilingual conversation" is a method or device for conducting a conversation with a user in multiple languages.

[1291] A "means for assessing user progress" is a method or device for measuring and analyzing a user's learning progress.

[1292] A "means for changing or adjusting a study plan" is a method or device for modifying an existing study plan based on a user's progress or needs.

[1293] "A means for balanced training of listening, speaking, reading, and writing skills" is a method or device for developing these language skills in a balanced manner.

[1294] The "means for analyzing the user's emotional state" refers to a method or device for analyzing the user's facial expression, tone of voice, etc. to determine their emotions.

[1295] A "means for dynamically adjusting a study plan" is a method or apparatus for optimizing a study plan in real time according to a user's emotional state and progress.

[1296] An "information processing device" is a computer system that performs multilingual conversations, assessment of learning progress, analysis of emotional states, and so on.

[1297] This invention relates to a system that generates a study plan based on user information, and dynamically adjusts the study plan by evaluating the user's progress and emotions while conducting multilingual conversations using a natural language processing engine and an emotion engine. This system is configured to effectively support the user's language learning.

[1298] The system mainly consists of a server and a terminal. The server receives user information and registers it in a database. User data includes name, age, language level, learning goals, etc. The server accesses the database and generates an individualized initial learning plan based on the user information. The terminal receives this learning plan and configures it to predict the user's emotional state. At this point, an emotional engine (e.g., Affectiva) is used to analyze the user's emotional state.

[1299] The learning plan is designed to balance the training of listening, speaking, reading, and writing skills. The device will speak to the user according to a set schedule, for example at 9:00 AM, asking, "Good morning! How did you sleep last night?" This dialogue is conducted using a natural language processing engine (e.g., GPT-3), which generates an appropriate response based on the user's input. If the user shows no interest, the emotion engine analyzes the user's facial expressions and tone of voice and changes the topic.

[1300] The device also offers listening, speaking, reading, and writing activities. For example, during listening, it plays a short story followed by questions about the story. When the user answers the questions, the device evaluates the answer and guides the user to the next step. The emotion engine analyzes the user's reaction and provides appropriate feedback.

[1301] The server periodically retrieves the user's learning progress data from the database and evaluates that progress. For example, if listening skills are improving but speaking skills are lacking, activities to strengthen speaking will be added to the next month's learning plan. The emotion engine also analyzes the user's motivation and optimizes the plan to increase activities that are likely to be rewarding. The server then sends the updated learning plan to the device, allowing for a smooth start to the next month's learning.

[1302] This system allows users to learn multiple languages ​​naturally in their daily lives, and by combining it with an emotion engine, it is possible to grasp the user's emotional state in real time and provide the optimal learning environment. Multilingual conversations utilizing generative AI models and prompt sentences allow users to constantly have new learning experiences. By utilizing technologies such as "generative AI models and prompt sentences," the system provides the most effective learning method for users.

[1303] Examples:

[1304] Example prompt: "Good morning! How did you sleep last night?" The next question or comment is automatically generated based on the user's response.

[1305] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1306] System program processing steps

[1307] Step 1: Entering user data and registering

[1308] 1-1: The user opens the new registration screen.

[1309] Input: The user enters information such as name, age, language level, and learning goals.

[1310] 1-2: The server receives the information sent by the user.

[1311] Data processing: Convert the received user information into database format.

[1312] 1-3: The server registers the information in the database.

[1313] Output: Generates a message confirming successful registration and sends it to the terminal.

[1314] What it does: The server uses a database management system (e.g., MySQL) to store data.

[1315] Step 2: Initial setup and customization

[1316] 2-1: The server reads the user information in the database.

[1317] Input: User information retrieved from the database.

[1318] 2-2: The server analyzes the user information and generates an initial learning plan.

[1319] Data processing: Generate a customized learning plan based on the user's age, language level, and learning goals.

[1320] 2-3: The device is configured to predict the user's emotional state based on the learning plan received from the server.

[1321] Input: The learning plan sent by the server.

[1322] 2-4: The device presents the learning plan to the user and asks for confirmation.

[1323] Output: Generates a preview of the learning plan and displays it to the user.

[1324] Specific operation: The device configures the emotion engine (e.g., Affectiva) and prepares it to analyze the user's emotional state.

[1325] Step 3: Multilingual conversation function

[1326] 3-1: The device starts a conversation session according to the set schedule.

[1327] Input: Schedule information based on your study plan.

[1328] 3-2: The user talks to the terminal.

[1329] Input: User voice or text input.

[1330] 3-3: The terminal analyzes the user's input using a natural language processing engine.

[1331] Data Computing: Using a natural language processing engine (e.g., GPT-3) to analyze user input and generate appropriate responses.

[1332] 3-4: The device uses an emotion engine to analyze the user's emotional state and dynamically adjust the content of the conversation.

[1333] Input: User's facial expressions and tone of voice.

[1334] Output: Generates tailored conversation content and presents it to the user.

[1335] Specific operation: The device responds to the user with the conversation content generated by the device via voice or text.

[1336] Step 4: Four English Skills Approach

[1337] 4-1: The device starts the learning activity at the set time.

[1338] Input: Activity schedule based on the lesson plan.

[1339] 4-2: The devices provide listening, speaking, reading, and writing activities.

[1340] Input: User actions or responses.

[1341] Data processing: Record and analyze user responses to each activity.

[1342] 4-3: The device uses an emotion engine to evaluate the user's reaction and provide feedback as needed.

[1343] Input: The user's emotional state.

[1344] Output: Generate feedback based on the evaluation results.

[1345] What it does: The listening activity plays a short story, and the speaking activity teaches the user pronunciation.

[1346] Step 5: Mother Tongue Acquisition Program

[1347] 5-1: The server periodically retrieves the user's learning progress data from the database.

[1348] Input: Progress data accumulated in the database.

[1349] 5-2: The server analyzes the progress data and updates the learning plan.

[1350] Data calculations: Analyze progress data to identify skills that users need to improve.

[1351] 5-3: The server uses the emotion engine to evaluate the user's emotional state and optimize the learning plan.

[1352] Input: User sentiment analysis data.

[1353] Output: Generates an optimized learning plan and sends it to the device.

[1354] 5-4: The server sends the updated learning plan to the device.

[1355] Output: Sends the updated study plan to the device, ready to start studying for the next month.

[1356] What it does: The server uses analytics tools to evaluate progress data and generate a learning plan with appropriate changes.

[1357] (Application example 2)

[1358] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1359] Conventional multilingual conversation systems lack the ability to dynamically analyze and respond to users' progress and emotional state in real time, resulting in insufficient learning efficiency and user interaction. Furthermore, they lacked the information processing capabilities to provide personalized responses when dealing with customers in physical stores. This can result in a lack of motivation to learn and a decline in user satisfaction. In particular, flexible responses that take into account the user's emotional state are essential for complex multilingual interactions.

[1360] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1361] In this invention, the server includes: means for inputting user information and generating a study plan based on the information; means for conducting multilingual conversations using a natural language processing engine based on the generated study plan; means for evaluating the user's progress and changing or adjusting the study plan based on the evaluation; means for providing activities for practicing listening, speaking, reading, and writing skills based on the generated study plan; means including an emotion analysis engine for analyzing the user's emotional state and providing an optimal response in real time; and means including an information processing device for responding to customers in a store. This makes it possible to analyze the user's study progress and emotional state in real time and provide an optimal study plan based on the analysis, while also enabling personalized and flexible customer service in physical stores.

[1362] "User" refers to an individual who uses the system.

[1363] "Information" refers to data about the user such as name, age, language level, learning goals, etc.

[1364] "Study Plan" refers to a schedule and content of study activities customized based on user information.

[1365] A "natural language processing engine" refers to a software component that analyzes text data and conducts conversations in multiple languages.

[1366] "Multilingual conversation" refers to interaction using multiple languages, during which real-time translation and dialogue control are carried out.

[1367] "Progress" refers to the results and achievement of a user's learning activities.

[1368] "Listening" refers to the skill of understanding spoken information.

[1369] "Speaking" refers to the skill of conveying information orally.

[1370] "Reading" refers to the skill of reading and understanding written information.

[1371] "Writing" refers to the skill of describing information using written text.

[1372] "Emotional state" refers to a user's current mental and emotional state.

[1373] An "emotion analysis engine" refers to a software component that analyzes a user's emotional state and responds appropriately.

[1374] "Information processing device" refers to a device that includes hardware and software for inputting, processing, and outputting data.

[1375] "Store" means a physical location for conducting sales activities.

[1376] This system starts by inputting user information and generating a learning plan based on that information. Specifically, the server receives information such as the user's name, age, language level, and learning goals, and registers it in a pre-prepared database. This forms the basis for customizing an appropriate learning plan for each user.

[1377] Next, the device uses a natural language processing engine to conduct a multilingual conversation based on the generated learning plan. Specifically, a schedule is created based on the information entered by the user, and the conversation proceeds according to that schedule. During the conversation, an emotion analysis engine analyzes the user's emotional state in real time, and the content of the conversation is adjusted accordingly.

[1378] The system constantly evaluates the user's progress and changes or adjusts the learning plan based on the evaluation results. This method makes it possible to provide an optimal learning plan that is always up to date. For example, if the user's listening skills improve, new skill practice will be incorporated to advance to the next level.

[1379] Furthermore, the system provides activities to practice listening, speaking, reading, and writing skills in a balanced manner. During the activities, the system analyzes the user's emotional state and provides appropriate feedback. This analysis is performed using software components such as EmotionAnalyzer and NLPModel.

[1380] A specific example is an information processing device installed in a cash register that interacts with the user in real time. For example, a smartphone or information robot can provide personalized guidance based on the user's name and learning goals. If the user appears unhappy, an emotion analysis engine can analyze their emotional state and flexibly adapt the response to improve user satisfaction.

[1381] An example of a prompt is as follows:

[1382] "Customer ID: 1

[1383] Customer Name: Sato

[1384] Language: Japanese

[1385] Age: 30

[1386] Emotional state: happy

[1387] Response: Personalized greetings and product recommendations

[1388] Output: 'Hello, Sato-san. How can I help you today? We have great deals on electronics today.'

[1389] Based on this prompt, the generative AI model generates appropriate answers and responses, making customer service in physical stores much smoother.

[1390] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1391] Step 1:

[1392] The server inputs user information and registers it in a database. The input information includes name, age, language level, learning goals, etc. Specifically, when a user inputs information using a smartphone or interface device, the server receives it and stores it in a database.

[1393] Input: User's name, age, language level, learning goal

[1394] Output: User information registered in the database

[1395] Step 2:

[1396] The server generates a learning plan based on the information stored in the database, including content tailored to the user's age, language level, and learning goals, and uses a sentiment analysis engine to customize the plan based on the user's emotional state.

[1397] Input: User information registered in the database

[1398] Output: A customized study plan

[1399] Step 3:

[1400] The device initiates a multilingual conversation using a natural language processing engine based on the generated learning plan. Specifically, the device proceeds with a dialogue with the user according to the learning plan and generates appropriate responses in real time using a generative AI model.

[1401] Input: Customized Study Plan

[1402] Output: User interaction, real-time generated responses

[1403] Step 4:

[1404] The server periodically evaluates the user's progress and modifies or adjusts the learning plan based on the evaluation, which includes activity log data and user feedback.

[1405] Input: Learning activity log data, user feedback

[1406] Output: Adjusted study plan

[1407] Step 5:

[1408] The device performs activities to practice listening, speaking, reading and writing skills based on a tailored learning plan, and during each activity, an emotion analysis engine analyzes the user's emotional state and provides appropriate feedback.

[1409] Input: Adjusted Study Plan

[1410] Output: User's emotional state data, feedback on activities

[1411] Step 6:

[1412] The server responds to users using the in-store information processing device. Specifically, when a user makes an inquiry to the in-store information processing device, the emotion analysis engine analyzes the user's emotional state based on the information and generates an appropriate response. Based on example prompt sentences, the generative AI model provides an appropriate response to the inquiry.

[1413] Input: In-store inquiry details, user's emotional state

[1414] Output: Personalized query response

[1415] Through the above processing steps, the system analyzes the user's learning progress and emotional state in real time, providing an optimal learning plan, and enabling personalized and flexible customer service in physical stores.

[1416] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1417] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1418] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1419] [Fourth embodiment]

[1420] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1421] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1422] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1423] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1424] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1425] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1426] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1427] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1428] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1429] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1430] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1431] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1432] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1433] The system of the present invention allows users to input their information, generate an individually customized learning plan based on that information, and engage in multilingual conversations using a generation AI. It also evaluates progress and adjusts the learning plan based on that evaluation, enabling the development of listening, speaking, reading, and writing skills in a balanced manner.

[1434] 1. Entering user data and registering

[1435] The server receives new user information and registers it in a database. User data includes name, age, language level, learning goals, etc.

[1436] Examples:

[1437] A user accesses the application and enters information such as name, age, language level, learning goals, etc. on the new registration screen. This information is received by the server and registered in the database.

[1438] 2. Initial setup and customization

[1439] The device generates an individual learning plan based on the user's information, which is customized according to the user's age, language level, and learning goals.

[1440] Examples:

[1441] When a 5-year-old user is enrolled, the device generates a learning plan for young children, including games, songs, and activities using basic English words and phrases.

[1442] 3. Multilingual conversation function

[1443] The device sets a schedule based on the learning plan and starts a multilingual conversation session according to that schedule. The conversation engine uses natural language processing to converse with the user in real time.

[1444] Examples:

[1445] At 9 a.m., the device speaks to the user, saying, "It's English time. Good morning! How did you sleep last night?" This dialogue is facilitated by generative AI, which automatically generates the next question or comment based on the user's response.

[1446] 4. Four English Skills Approach

[1447] The device offers a variety of activities to balance and develop users' listening, speaking, reading and writing skills.

[1448] Examples:

[1449] During the listening section, the device will play a short story to the user and then ask questions about the story, while during the speaking section, the device will have the user pronounce a simple phrase and evaluate their pronunciation to guide them to the next step.

[1450] 5. Mother Tongue Acquisition Program

[1451] The server periodically evaluates the user's learning progress and analyzes the progress data stored in the database, and based on this evaluation, modifies or adjusts the learning plan.

[1452] Examples:

[1453] At the end of the month, during progress assessment, the server determines that the user's listening skills have improved, but their speaking skills are lacking. Based on this, speaking-focused activities are added to the next month's learning plan.

[1454] The system of the present invention naturally integrates multilingual exposure into the user's living environment, allowing the user to naturally acquire a second language in their daily life and grow as a bilingual. This system is particularly effective in language acquisition from an early age, and solves the problems facing the modern education system.

[1455] The processing flow will be explained below.

[1456] Processing flow

[1457] Entering user data and registering

[1458] Step 1:

[1459] Users access the application and enter information such as their name, age, language level, and learning goals on the new registration screen.

[1460] Step 2:

[1461] The terminal receives the entered user information and sends it to the server.

[1462] Step 3:

[1463] The server registers the received user information in a database and generates a user ID.

[1464] First-time setup and customization

[1465] Step 1:

[1466] The server retrieves user information from a database and generates an initial learning plan based on this.

[1467] Step 2:

[1468] The device presents the generated learning plan to the user and asks for confirmation of the learning plan.

[1469] Step 3:

[1470] The user reviews the proposed study plan and requests adjustments if necessary.

[1471] Multilingual conversation function

[1472] Step 1:

[1473] The device starts a conversation session on a specified schedule based on the learning plan.

[1474] Step 2:

[1475] MimiTomo (robot) uses a natural language processing engine to converse with users in the selected language.

[1476] Step 3:

[1477] Users have the opportunity to converse with the robot and use language naturally.

[1478] Four English Skills Approach

[1479] Step 1:

[1480] The device generates and provides activities for practicing listening, speaking, reading and writing skills.

[1481] Step 2:

[1482] The user performs each skill practice activity provided by the terminal.

[1483] Step 3:

[1484] The terminal collects the user's activity results and sends them to the server.

[1485] Mother Tongue Acquisition Program

[1486] Step 1:

[1487] The server periodically evaluates the user's learning progress and updates the database based on the evaluation results.

[1488] Step 2:

[1489] The server analyzes the progress data and adjusts the learning plan as needed.

[1490] Step 3:

[1491] The device will present the user with a new, adjusted learning plan and begin the next learning cycle.

[1492] In this way, the system generates an individualized learning plan based on the user's information, engages in daily multilingual conversations based on that plan, and provides activities to train the listening, speaking, reading, and writing skills in a balanced manner.By continuously evaluating the user's progress, the system supports effective bilingual development.

[1493] Example 1

[1494] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1495] In multilingual learning, there is a need for systems that can take into account individual differences, provide users with a learning plan that is tailored to their needs, and adjust it as needed based on their progress. In particular, there is a lack of effective methods for balanced development of listening, speaking, reading, and writing skills. Another issue is the difficulty of effectively incorporating real-time multilingual conversation functions.

[1496] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1497] In this invention, the server includes means for inputting user information and generating a study plan based on the information, means for conducting multilingual conversations using a natural language processing engine based on the generated study plan, means for evaluating the user's progress and changing or adjusting the study plan based on the evaluation, and means for providing activities for practicing the skills of auditory comprehension, speaking, reading comprehension, and writing based on the generated study plan. This makes it possible to create and adjust a study plan suited to the user and develop each skill in a balanced manner.

[1498] "User" refers to an individual who uses the system to acquire multiple languages.

[1499] "Information" refers to data about the user such as name, age, language level, learning goals, etc.

[1500] "Study Plan" refers to a customized study schedule and content based on user information.

[1501] A "natural language processing engine" refers to technology that enables computers to understand and generate human language.

[1502] A "multilingual conversation" refers to a dialogue that takes place using more than one language.

[1503] "Progress" refers to the state of how far a user has progressed or improved on their learning plan.

[1504] "Evaluation" refers to the process of measuring a user's learning progress and analyzing the results.

[1505] "Auditory comprehension" refers to the skill of understanding language heard.

[1506] "Speech" refers to the skill of speaking a language.

[1507] "Reading comprehension" refers to the skill of understanding written language.

[1508] "Writing" refers to the skill of using language to write sentences.

[1509] "Activity" refers to specific exercises or tasks that a user performs to practice each skill.

[1510] The system of the present invention inputs user information, generates an individually customized learning plan based on that information, and conducts multilingual conversations using a generative AI model. It also evaluates progress and adjusts the learning plan based on that evaluation, enabling the development of auditory comprehension, speaking, reading comprehension, and writing skills in a balanced manner.

[1511] 1. Entering user data and registering

[1512] The server receives new user information and registers it in a database. User data includes name, age, language level, learning goals, etc.

[1513] Examples:

[1514] A user accesses the application and enters information such as name, age, language level, learning goals, etc. on the new registration screen. This information is received by the server and registered in the database.

[1515] 2. Initial setup and customization

[1516] The device generates an individual learning plan based on the user's information, which is customized according to the user's age, language level, and learning goals.

[1517] Examples:

[1518] If a 5-year-old user is enrolled, the device will generate a learning plan for young children, including games, songs, and activities using basic English words and phrases.

[1519] 3. Multilingual conversation function

[1520] The device sets a schedule based on the study plan and starts a multilingual conversation session according to that schedule. The conversation engine uses natural language processing to converse with the user in real time.

[1521] Examples:

[1522] At 9 a.m., the device speaks to the user, "It's time for English. Good morning! How did you sleep last night?" This dialogue is driven by a generative AI model, which automatically generates the next question or comment based on the user's response.

[1523] 4. Four English Skills Approach

[1524] The device offers a variety of activities to balance the user's auditory comprehension, speaking, reading, and writing skills.

[1525] Examples:

[1526] During the listening section, the device will play a short story to the user and then ask questions about the story. During the speaking section, the device will have the user pronounce simple phrases and evaluate their pronunciation to guide them to the next step.

[1527] 5. Mother Tongue Acquisition Program

[1528] The server periodically evaluates the user's learning progress and analyzes the progress data stored in the database, and based on this evaluation, modifies or adjusts the learning plan.

[1529] Examples:

[1530] At the end of the month, the server assesses the progress and finds that the user's listening skills have improved, but their speaking skills are lacking. Based on this, activities to strengthen speaking are added to the next month's lesson plan.

[1531] Through this entire system, multilingual exposure can be naturally incorporated into the user's living environment, allowing them to naturally acquire a second language in their daily lives and grow as bilinguals. This is particularly effective for language acquisition from an early age, and solves the problems facing the modern education system.

[1532] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1533] Step 1:

[1534] Entering user data and registering

[1535] Input: The user enters information such as name, age, language level, and learning goals on the application's sign-up screen.

[1536] Specific behavior: The data entered by the user is sent to the server through the input form.

[1537] Data processing: The server formats the received user data into an appropriate format (e.g., converts it to JSON format).

[1538] Output: Correctly formatted user data.

[1539] Step 2:

[1540] Receiving information and registering in the database

[1541] Input: The server receives the formatted user data.

[1542] What happens: The server checks the format of the data and issues the appropriate SQL query to store it in the database.

[1543] Data manipulation: Insert user data into the database using SQL queries.

[1544] Output: User data is registered in the database.

[1545] Step 3:

[1546] User information acquisition

[1547] Input: User data registered on the server.

[1548] Specific operation: The device obtains user data from the server through an API call.

[1549] Data processing: Converting received data into internal data structures.

[1550] Output: User data stored in the device.

[1551] Step 4:

[1552] Generate a lesson plan

[1553] Input: User data stored on the device.

[1554] What it does: The lesson plan generator generates an appropriate lesson plan based on the user's age, language level, and learning goals.

[1555] Data processing: The study plan generator analyzes user data and creates a customized study plan that is optimal for them.

[1556] Output: A customized learning plan for the user.

[1557] Step 5:

[1558] Scheduling

[1559] Input: Schedule information based on your study plan.

[1560] Specific behavior: The device uses the calendar API to automatically reserve a session for the specified date and time.

[1561] Data processing: Convert schedule information into calendar format and register it on the user's device.

[1562] Output: The configured schedule.

[1563] Step 6:

[1564] Starting a multilingual conversation session

[1565] Input: A prompt statement based on the lesson plan.

[1566] Specific operation: The device sends a prompt to the generative AI model and begins a dialogue with the user in real time.

[1567] Data processing: Input the prompt sentence into the natural language processing engine and receive a response from the generative AI model.

[1568] Output: An ongoing multilingual conversation with the user.

[1569] Step 7:

[1570] Skills Assessment

[1571] Input: User behavior data during a multilingual conversation session.

[1572] How it works: The device uses speech recognition technology and natural language processing to evaluate the user's pronunciation and responses.

[1573] Data processing: Analyze the collected data and generate metrics to assess the user's skill level.

[1574] Output: User skill assessment results.

[1575] Step 8:

[1576] Collecting and evaluating progress data

[1577] Input: Skill assessment results after each session.

[1578] Specific operation: The server collects and analyzes the user's learning progress data sent from the device.

[1579] Data processing: The collected data is input into a statistical model to evaluate progress.

[1580] Output: Progress assessment results and reports.

[1581] Step 9:

[1582] Adjusting your study plan

[1583] Input: Progress assessment results.

[1584] Specific operation: The server generates a new learning plan based on the progress assessment and distributes it to the device.

[1585] Data processing: Applying algorithms to progress data to create new, customized learning plans.

[1586] Output: A tailored learning plan.

[1587] (Application example 1)

[1588] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1589] Increasing the efficiency of multilingual learning is a challenge facing the modern education system. Furthermore, with limited options for effective language learning in busy daily lives, there is a need for learning methods that allow for effective use of time, especially while traveling. There is also a need for learning plans tailored to the age and language level of students, from children to adults, that foster a balanced development of listening, speaking, reading, and writing skills. To address these challenges, a system is needed that can provide customized learning plans and progress management tailored to each user's individual needs.

[1590] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1591] The invention includes means for inputting user information and generating a learning plan based on the information, means for conducting multilingual conversations using a natural language processing engine based on the generated learning plan, means for evaluating the user's progress and changing or adjusting the learning plan based on the evaluation, means for providing activities to practice listening, speaking, reading, and writing skills based on the generated learning plan, and means for providing multilingual learning using a display, microphone, and speaker in the vehicle.

[1592] This allows passengers to effectively utilize their time while traveling in a vehicle to study multiple languages, and enables them to efficiently develop their language skills in a balanced manner through a learning plan customized for each user.In addition, real-time dialogue and progress management can maximize the effectiveness of their learning.

[1593] "User Information" refers to personal data necessary to generate a learning plan, such as the learner's name, age, language level, and learning goals.

[1594] "Study Plan" refers to specific learning content and schedules that are individually customized based on the user's information and designed to develop listening, speaking, reading, and writing skills in a balanced manner.

[1595] A "natural language processing engine" refers to a software system that uses generative AI to conduct natural conversations in multiple languages.

[1596] "Progress assessment" refers to the process of quantitatively and qualitatively evaluating the results of the learning activities that a user has carried out based on their learning plan.

[1597] "Listening, speaking, reading and writing skills" refers to the four basic skill areas in language learning, with corresponding practice activities.

[1598] "In-vehicle display" refers to the screen installed inside an autonomous vehicle that displays learning content and provides interactive instructions.

[1599] A "microphone" refers to an input device that picks up a user's voice and analyzes the voice in real time.

[1600] "Speaker" refers to an audio device that allows interaction between the user and the system through audio output.

[1601] A "robot device" refers to a hardware device used to interact with a user in real time, and has functions such as voice recognition and voice synthesis.

[1602] This invention is a system for providing multilingual learning within an autonomous vehicle, which has the function of generating a learning plan based on user information and conducting multilingual conversations in real time based on that plan. The system includes the following elements:

[1603] 1. Entering user data and registering

[1604] The server receives information such as the user's name, age, language level, and learning goals provided by the user via a display, smartphone, or other device when they board an autonomous vehicle, and registers this information in a database, creating a foundation for generating learning plans tailored to the user's individual needs.

[1605] 2. Create and customize your study plan

[1606] The server uses a generative AI model based on registered user information to generate individual learning plans. For example, it can provide a business English course for adults commuting to work, and a playful basic English learning plan for children. The learning plans include a balanced mix of listening, speaking, reading, and writing skills.

[1607] 3. Multilingual conversation function

[1608] The device sets a schedule based on the learning plan and starts multilingual conversation sessions according to that schedule. Using the vehicle's microphone and speaker, the device can communicate with the user in real time using a natural language processing engine, allowing users to efficiently improve their language skills while on the move.

[1609] 4. Four Skills Approach

[1610] The device offers multifaceted activities to practice listening, speaking, reading, and writing skills based on a learning plan. For example, during listening, the device plays a short story followed by questions to check comprehension. During speaking, the device evaluates the user's pronunciation in real time and provides feedback.

[1611] 5. Progress assessment and plan adjustment

[1612] The server periodically evaluates the user's progress and analyzes the progress data stored in the database. Based on this evaluation, the server changes or adjusts the learning plan to help the user improve their skills. For example, if listening skills are improving but speaking skills are lacking, the server will add activities to strengthen speaking in the next month.

[1613] Specific examples

[1614] A user boards an autonomous vehicle and registers by entering their name, age, language level, and learning goals on the display terminal. The server then generates a customized learning plan based on the user's information and initiates a multilingual conversation session through the vehicle's speakers and display.

[1615] Prompt Sentence Examples

[1616] "Enter the user's name, age, language level, and learning goals, and generate a program to charge them in the car's interface."

[1617] This system allows users to study multiple languages ​​according to an individually customized learning plan, making effective use of their time inside an autonomous vehicle. It is expected that real-time dialogue and progress management will maximize learning effectiveness.

[1618] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1619] Step 1:

[1620] When a user gets into an autonomous vehicle, they input information via a display or smartphone. This information includes their name, age, language level, and learning goals, and the input is sent to a server. The server receives this information and registers it in a database. The input of this step is user information, and the output is the user information registered in the database.

[1621] Step 2:

[1622] The server uses a generative AI model to generate individual study plans based on registered user information. The generated study plans are customized according to the user's age, language level, and learning goals. For example, an adult user aiming for business English would be provided with a study plan focused on business situations. The input for this step is user information registered in the database, and the output is a customized study plan.

[1623] Step 3:

[1624] The device sets a schedule based on the learning plan sent from the server. Specifically, it uses the vehicle's display and speaker to indicate the required activities to the user and start the learning session. For example, it sets an English listening session to start at 9:00 AM. The input of this step is the customized learning plan, and the output is the set schedule.

[1625] Step 4:

[1626] The device initiates a multilingual conversation session according to a set schedule. It collects user utterances through a microphone in the vehicle and analyzes them through a natural language processing engine. It uses a generative AI model on the analyzed data to generate an appropriate reply and respond to the user through the speaker. The input for this step is the user's utterance data, and the output is the generated response.

[1627] Step 5:

[1628] The device provides activities for practicing listening, speaking, reading, and writing skills. For example, a listening session plays a short story followed by questions to check comprehension. A speaking session allows the user to practice pronunciation and evaluate its accuracy with a speech recognition system. The input for this step is the user's activity data based on the learning plan, and the output is the evaluation result.

[1629] Step 6:

[1630] The server periodically evaluates the user's progress and analyzes the progress data stored in the database. Based on this evaluation, it modifies or adjusts the user's learning plan. For example, if the user's listening skills are improving but their speaking skills are lacking, it will add activities to strengthen their speaking skills next month. The input of this step is the progress data stored in the database, and the output is an adjusted learning plan.

[1631] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1632] The system generates a learning plan based on user information, and utilizes a natural language processing engine and an emotion engine to evaluate the user's progress and emotions while conducting multilingual conversations, dynamically adjusting the learning plan. By nurturing listening, speaking, reading, and writing skills in a balanced manner and taking into account the user's emotional state, the system provides an optimal learning experience.

[1633] 1. Entering user data and registering

[1634] The server receives new user information and registers it in a database. User data includes name, age, language level, learning goals, etc.

[1635] Examples:

[1636] A user accesses the application and enters information such as name, age, language level, learning goals, etc. on the new registration screen. This information is received by the server and registered in the database.

[1637] 2. Initial setup and customization

[1638] The device generates an individual learning plan based on the user's information, which is customized according to the user's age, language level, and learning goals. Furthermore, an emotion engine is used to configure the plan to take into account the user's emotional state.

[1639] Examples:

[1640] When a five-year-old user registers, the device generates a learning plan for young children, including games and songs, activities using basic English words and phrases, and selects content that will interest the user and analyzes it with an emotion engine.

[1641] 3. Multilingual conversation function

[1642] The device sets a schedule based on the learning plan and starts a multilingual conversation session according to that schedule. The conversation engine uses natural language processing to converse with the user in real time. Furthermore, the emotion engine analyzes the user's emotions and adjusts the content of the conversation accordingly.

[1643] Examples:

[1644] At 9 a.m., the device speaks to the user, saying, "It's time for English. Good morning! How did you sleep last night?" This dialogue is conducted by a generative AI, which automatically generates the next question or comment based on the user's response. If the user shows no interest, the emotion engine analyzes the user's facial expressions and tone of voice and changes the topic.

[1645] 4. Four English Skills Approach

[1646] The device generates and provides activities to balance listening, speaking, reading, and writing skills, and uses an emotion engine to analyze the user's emotional state and provide appropriate feedback.

[1647] Examples:

[1648] During the listening section, the device tells the user a short story and then asks them questions about it. During the speaking section, the device asks the user to pronounce simple phrases, and the device evaluates their pronunciation and guides them to the next step. An emotion engine analyzes whether the user is enjoying themselves or struggling and adjusts its approach accordingly.

[1649] 5. Mother Tongue Acquisition Program

[1650] The server periodically evaluates the user's learning progress and analyzes the progress data stored in the database. Based on this evaluation, the server changes or adjusts the learning plan. It also uses an emotion engine to consider the user's emotional state and provide an optimal plan that maintains the user's motivation to learn.

[1651] Examples:

[1652] At the end of the month, the server will assess the user's progress and determine that their listening skills have improved, but their speaking skills are lacking. Based on this, activities to strengthen speaking will be added to the next month's learning plan. Furthermore, the emotion engine will analyze the user's motivation and optimize the learning plan to increase activities that they find challenging.

[1653] This system allows users to naturally acquire a second language in their daily lives and grow as bilinguals. By combining it with an emotion engine, it is possible to grasp the user's emotional state in real time and provide an optimal learning environment. As a result, an environment is created where multilingual conversations are constantly taking place, allowing users to grow naturally in a bilingual environment.

[1654] The processing flow will be explained below.

[1655] Processing flow

[1656] Entering user data and registering

[1657] Step 1:

[1658] Users access the application and enter information such as their name, age, language level, and learning goals on the new registration screen.

[1659] Step 2:

[1660] The terminal receives the entered user information and sends it to the server.

[1661] Step 3:

[1662] The server registers the received user information in a database and generates a user ID.

[1663] First-time setup and customization

[1664] Step 1:

[1665] The server retrieves user information from a database and generates an initial learning plan based on this.

[1666] Step 2:

[1667] The device presents the generated learning plan to the user and asks for confirmation of the learning plan.

[1668] Step 3:

[1669] The user reviews the proposed study plan and requests adjustments if necessary.

[1670] Step 4:

[1671] The server adjusts the learning plan based on user feedback and determines the final plan.

[1672] Multilingual conversation function

[1673] Step 1:

[1674] The device starts a conversation session on a specified schedule based on the learning plan.

[1675] Step 2:

[1676] MimiTomo (robot) uses a natural language processing engine to converse with users in the selected language.

[1677] Step 3:

[1678] The emotion engine analyzes the user's facial expressions and tone of voice to determine the user's emotional state.

[1679] Step 4:

[1680] MimiTomo (the robot) adjusts the conversation based on the user's emotional state, for example, changing the topic if the user is losing interest.

[1681] Step 5:

[1682] Users have the opportunity to converse with the robot and use language naturally.

[1683] Four English Skills Approach

[1684] Step 1:

[1685] The device generates and provides activities for practicing listening, speaking, reading and writing skills.

[1686] Step 2:

[1687] The user performs each skill practice activity provided by the terminal.

[1688] Step 3:

[1689] The emotion engine analyzes the user's emotional state during an activity and adjusts the difficulty and content of the activity if motivation is declining.

[1690] Step 4:

[1691] The terminal collects the user's activity results and sends them to the server.

[1692] Mother Tongue Acquisition Program

[1693] Step 1:

[1694] The server periodically evaluates the user's learning progress and updates the database based on the evaluation results.

[1695] Step 2:

[1696] The server analyzes the progress data and adjusts the learning plan as needed.

[1697] Step 3:

[1698] The emotion engine also evaluates the user's emotional data and reflects this in adjusting the learning plan.

[1699] Step 4:

[1700] The device will present the user with a new, adjusted learning plan and begin the next learning cycle.

[1701] Specific examples

[1702] Entering user data and registering

[1703] Step 1:

[1704] A user accesses the application and registers by entering their name "Taro," age "5 years old," language level "Beginner," and learning goal "Bilingual."

[1705] Step 2:

[1706] The terminal receives the input information and transmits it to the server.

[1707] Step 3:

[1708] The server registers the received information in the database and generates a user ID "001."

[1709] First-time setup and customization

[1710] Step 1:

[1711] The server generates an individual learning plan, "English Learning Plan for Toddlers," based on the user ID "001."

[1712] Step 2:

[1713] The device will present the generated learning plan to the user and ask them to "confirm."

[1714] Step 3:

[1715] The user reviews the plan and requests any necessary adjustments (e.g., adding specific activities).

[1716] Step 4:

[1717] The server adjusts the learning plan based on the request and finalizes the plan.

[1718] Multilingual conversation function

[1719] Step 1:

[1720] The device will set up and start an English conversation session at 9am every morning.

[1721] Step 2:

[1722] MimiTomo (robot) says, "Good morning, Taro! How did you sleep last night?"

[1723] Step 3:

[1724] The emotion engine analyzes the user's facial expression and determines that they are "showing interest."

[1725] Step 4:

[1726] MimiTomo (robot) continues the topic by saying, "Let's talk about your favorite toy."

[1727] Step 5:

[1728] The user responds in English, "My favorite toy is a car."

[1729] Four English Skills Approach

[1730] Step 1:

[1731] The device generates a listening activity and plays a short story.

[1732] Step 2:

[1733] Users listen to a story and answer related questions.

[1734] Step 3:

[1735] The emotion engine analyzes the user's level of concentration and determines that they are "concentrated."

[1736] Step 4:

[1737] The device provides additional activity questions to promote deeper understanding.

[1738] Mother Tongue Acquisition Program

[1739] Step 1:

[1740] The server evaluates the user's progress at the end of each month and updates the database.

[1741] Step 2:

[1742] The server analyzes the progress data and determines that "speaking skills need to be improved."

[1743] Step 3:

[1744] The emotion engine analyzes the user's motivation data and selects activities that they find rewarding.

[1745] Step 4:

[1746] The device will present the user with a new, adjusted learning plan and begin the next learning cycle.

[1747] Example 2

[1748] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1749] Conventional educational systems have had difficulty dynamically adjusting learning plans based on each learner's progress and emotional state. Especially in multilingual conversation learning, there has been a lack of appropriate feedback and activities that take into account the learner's emotional state. As a result, it has been difficult to maintain learner motivation and acquire skills efficiently.

[1750] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1751] In this invention, the server includes means for inputting user information and generating a study plan based on the information, means for conducting multilingual conversations using a natural language processing engine based on the generated study plan, means for evaluating the user's progress and changing or adjusting the study plan based on the evaluation, means for training the listening, speaking, reading, and writing skills in a balanced manner, and means for analyzing the user's emotional state and dynamically adjusting the study plan, thereby making it possible to provide an optimal learning experience according to the progress and emotional state of each individual learner.

[1752] "Means for inputting user information" refers to a method or device by which a user inputs information such as name, age, language level, learning goals, etc. into the system.

[1753] A "means for generating a study plan" is a method or device for creating an individualized study plan based on user information.

[1754] A "natural language processing engine" is a technology that uses artificial intelligence to analyze user input and generate appropriate responses.

[1755] A "means for conducting a multilingual conversation" is a method or device for conducting a conversation with a user in multiple languages.

[1756] A "means for assessing user progress" is a method or device for measuring and analyzing a user's learning progress.

[1757] A "means for changing or adjusting a study plan" is a method or device for modifying an existing study plan based on a user's progress or needs.

[1758] "A means for balanced training of listening, speaking, reading, and writing skills" is a method or device for developing these language skills in a balanced manner.

[1759] The "means for analyzing the user's emotional state" refers to a method or device for analyzing the user's facial expression, tone of voice, etc. to determine their emotions.

[1760] A "means for dynamically adjusting a study plan" is a method or apparatus for optimizing a study plan in real time according to a user's emotional state and progress.

[1761] An "information processing device" is a computer system that performs multilingual conversations, assessment of learning progress, analysis of emotional states, and so on.

[1762] This invention relates to a system that generates a study plan based on user information, and dynamically adjusts the study plan by evaluating the user's progress and emotions while conducting multilingual conversations using a natural language processing engine and an emotion engine. This system is configured to effectively support the user's language learning.

[1763] The system mainly consists of a server and a terminal. The server receives user information and registers it in a database. User data includes name, age, language level, learning goals, etc. The server accesses the database and generates an individualized initial learning plan based on the user information. The terminal receives this learning plan and configures it to predict the user's emotional state. At this point, an emotional engine (e.g., Affectiva) is used to analyze the user's emotional state.

[1764] The learning plan is designed to balance the training of listening, speaking, reading, and writing skills. The device will speak to the user according to a set schedule, for example at 9:00 AM, asking, "Good morning! How did you sleep last night?" This dialogue is conducted using a natural language processing engine (e.g., GPT-3), which generates an appropriate response based on the user's input. If the user shows no interest, the emotion engine analyzes the user's facial expressions and tone of voice and changes the topic.

[1765] The device also offers listening, speaking, reading, and writing activities. For example, during listening, it plays a short story followed by questions about the story. When the user answers the questions, the device evaluates the answer and guides the user to the next step. The emotion engine analyzes the user's reaction and provides appropriate feedback.

[1766] The server periodically retrieves the user's learning progress data from the database and evaluates that progress. For example, if listening skills are improving but speaking skills are lacking, activities to strengthen speaking will be added to the next month's learning plan. The emotion engine also analyzes the user's motivation and optimizes the plan to increase activities that are likely to be rewarding. The server then sends the updated learning plan to the device, allowing for a smooth start to the next month's learning.

[1767] This system allows users to learn multiple languages ​​naturally in their daily lives, and by combining it with an emotion engine, it is possible to grasp the user's emotional state in real time and provide the optimal learning environment. Multilingual conversations utilizing generative AI models and prompt sentences allow users to constantly have new learning experiences. By utilizing technologies such as "generative AI models and prompt sentences," the system provides the most effective learning method for users.

[1768] Examples:

[1769] Example prompt: "Good morning! How did you sleep last night?" The next question or comment is automatically generated based on the user's response.

[1770] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1771] System program processing steps

[1772] Step 1: Entering user data and registering

[1773] 1-1: The user opens the new registration screen.

[1774] Input: The user enters information such as name, age, language level, and learning goals.

[1775] 1-2: The server receives the information sent by the user.

[1776] Data processing: Convert the received user information into database format.

[1777] 1-3: The server registers the information in the database.

[1778] Output: Generates a message confirming successful registration and sends it to the terminal.

[1779] What it does: The server uses a database management system (e.g., MySQL) to store data.

[1780] Step 2: Initial setup and customization

[1781] 2-1: The server reads the user information in the database.

[1782] Input: User information retrieved from the database.

[1783] 2-2: The server analyzes the user information and generates an initial learning plan.

[1784] Data processing: Generate a customized learning plan based on the user's age, language level, and learning goals.

[1785] 2-3: The device is configured to predict the user's emotional state based on the learning plan received from the server.

[1786] Input: The learning plan sent by the server.

[1787] 2-4: The device presents the learning plan to the user and asks for confirmation.

[1788] Output: Generates a preview of the learning plan and displays it to the user.

[1789] Specific operation: The device configures the emotion engine (e.g., Affectiva) and prepares it to analyze the user's emotional state.

[1790] Step 3: Multilingual conversation function

[1791] 3-1: The device starts a conversation session according to the set schedule.

[1792] Input: Schedule information based on your study plan.

[1793] 3-2: The user talks to the terminal.

[1794] Input: User voice or text input.

[1795] 3-3: The terminal analyzes the user's input using a natural language processing engine.

[1796] Data Computing: Using a natural language processing engine (e.g., GPT-3) to analyze user input and generate appropriate responses.

[1797] 3-4: The device uses an emotion engine to analyze the user's emotional state and dynamically adjust the content of the conversation.

[1798] Input: User's facial expressions and tone of voice.

[1799] Output: Generates tailored conversation content and presents it to the user.

[1800] Specific operation: The device responds to the user with the conversation content generated by the device via voice or text.

[1801] Step 4: Four English Skills Approach

[1802] 4-1: The device starts the learning activity at the set time.

[1803] Input: Activity schedule based on the lesson plan.

[1804] 4-2: The devices provide listening, speaking, reading, and writing activities.

[1805] Input: User actions or responses.

[1806] Data processing: Record and analyze user responses to each activity.

[1807] 4-3: The device uses an emotion engine to evaluate the user's reaction and provide feedback as needed.

[1808] Input: The user's emotional state.

[1809] Output: Generate feedback based on the evaluation results.

[1810] What it does: The listening activity plays a short story, and the speaking activity teaches the user pronunciation.

[1811] Step 5: Mother Tongue Acquisition Program

[1812] 5-1: The server periodically retrieves the user's learning progress data from the database.

[1813] Input: Progress data accumulated in the database.

[1814] 5-2: The server analyzes the progress data and updates the learning plan.

[1815] Data calculations: Analyze progress data to identify skills that users need to improve.

[1816] 5-3: The server uses the emotion engine to evaluate the user's emotional state and optimize the learning plan.

[1817] Input: User sentiment analysis data.

[1818] Output: Generates an optimized learning plan and sends it to the device.

[1819] 5-4: The server sends the updated learning plan to the device.

[1820] Output: Sends the updated study plan to the device, ready to start studying for the next month.

[1821] What it does: The server uses analytics tools to evaluate progress data and generate a learning plan with appropriate changes.

[1822] (Application example 2)

[1823] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1824] Conventional multilingual conversation systems lack the ability to dynamically analyze and respond to users' progress and emotional state in real time, resulting in insufficient learning efficiency and user interaction. Furthermore, they lacked the information processing capabilities to provide personalized responses when dealing with customers in physical stores. This can result in a lack of motivation to learn and a decline in user satisfaction. In particular, flexible responses that take into account the user's emotional state are essential for complex multilingual interactions.

[1825] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1826] In this invention, the server includes: means for inputting user information and generating a study plan based on the information; means for conducting multilingual conversations using a natural language processing engine based on the generated study plan; means for evaluating the user's progress and changing or adjusting the study plan based on the evaluation; means for providing activities for practicing listening, speaking, reading, and writing skills based on the generated study plan; means including an emotion analysis engine for analyzing the user's emotional state and providing an optimal response in real time; and means including an information processing device for responding to customers in a store. This makes it possible to analyze the user's study progress and emotional state in real time and provide an optimal study plan based on the analysis, while also enabling personalized and flexible customer service in physical stores.

[1827] "User" refers to an individual who uses the system.

[1828] "Information" refers to data about the user such as name, age, language level, learning goals, etc.

[1829] "Study Plan" refers to a schedule and content of study activities customized based on user information.

[1830] A "natural language processing engine" refers to a software component that analyzes text data and conducts conversations in multiple languages.

[1831] "Multilingual conversation" refers to interaction using multiple languages, during which real-time translation and dialogue control are carried out.

[1832] "Progress" refers to the results and achievement of a user's learning activities.

[1833] "Listening" refers to the skill of understanding spoken information.

[1834] "Speaking" refers to the skill of conveying information orally.

[1835] "Reading" refers to the skill of reading and understanding written information.

[1836] "Writing" refers to the skill of describing information using written text.

[1837] "Emotional state" refers to a user's current mental and emotional state.

[1838] An "emotion analysis engine" refers to a software component that analyzes a user's emotional state and responds appropriately.

[1839] "Information processing device" refers to a device that includes hardware and software for inputting, processing, and outputting data.

[1840] "Store" means a physical location for conducting sales activities.

[1841] This system starts by inputting user information and generating a learning plan based on that information. Specifically, the server receives information such as the user's name, age, language level, and learning goals, and registers it in a pre-prepared database. This forms the basis for customizing an appropriate learning plan for each user.

[1842] Next, the device uses a natural language processing engine to conduct a multilingual conversation based on the generated learning plan. Specifically, a schedule is created based on the information entered by the user, and the conversation proceeds according to that schedule. During the conversation, an emotion analysis engine analyzes the user's emotional state in real time, and the content of the conversation is adjusted accordingly.

[1843] The system constantly evaluates the user's progress and changes or adjusts the learning plan based on the evaluation results. This method makes it possible to provide an optimal learning plan that is always up to date. For example, if the user's listening skills improve, new skill practice will be incorporated to advance to the next level.

[1844] Furthermore, the system provides activities to practice listening, speaking, reading, and writing skills in a balanced manner. During the activities, the system analyzes the user's emotional state and provides appropriate feedback. This analysis is performed using software components such as EmotionAnalyzer and NLPModel.

[1845] A specific example is an information processing device installed in a cash register that interacts with the user in real time. For example, a smartphone or information robot can provide personalized guidance based on the user's name and learning goals. If the user appears unhappy, an emotion analysis engine can analyze their emotional state and flexibly adapt the response to improve user satisfaction.

[1846] An example of a prompt is as follows:

[1847] "Customer ID: 1

[1848] Customer Name: Sato

[1849] Language: Japanese

[1850] Age: 30

[1851] Emotional state: happy

[1852] Response: Personalized greetings and product recommendations

[1853] Output: 'Hello, Sato-san. How can I help you today? We have great deals on electronics today.'

[1854] Based on this prompt, the generative AI model generates appropriate answers and responses, making customer service in physical stores much smoother.

[1855] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1856] Step 1:

[1857] The server inputs user information and registers it in a database. The input information includes name, age, language level, learning goals, etc. Specifically, when a user inputs information using a smartphone or interface device, the server receives it and stores it in a database.

[1858] Input: User's name, age, language level, learning goal

[1859] Output: User information registered in the database

[1860] Step 2:

[1861] The server generates a learning plan based on the information stored in the database, including content tailored to the user's age, language level, and learning goals, and uses a sentiment analysis engine to customize the plan based on the user's emotional state.

[1862] Input: User information registered in the database

[1863] Output: A customized study plan

[1864] Step 3:

[1865] The device initiates a multilingual conversation using a natural language processing engine based on the generated learning plan. Specifically, the device proceeds with a dialogue with the user according to the learning plan and generates appropriate responses in real time using a generative AI model.

[1866] Input: Customized Study Plan

[1867] Output: User interaction, real-time generated responses

[1868] Step 4:

[1869] The server periodically evaluates the user's progress and modifies or adjusts the learning plan based on the evaluation, which includes activity log data and user feedback.

[1870] Input: Learning activity log data, user feedback

[1871] Output: Adjusted study plan

[1872] Step 5:

[1873] The device performs activities to practice listening, speaking, reading and writing skills based on a tailored learning plan, and during each activity, an emotion analysis engine analyzes the user's emotional state and provides appropriate feedback.

[1874] Input: Adjusted Study Plan

[1875] Output: User's emotional state data, feedback on activities

[1876] Step 6:

[1877] The server responds to users using the in-store information processing device. Specifically, when a user makes an inquiry to the in-store information processing device, the emotion analysis engine analyzes the user's emotional state based on the information and generates an appropriate response. Based on example prompt sentences, the generative AI model provides an appropriate response to the inquiry.

[1878] Input: In-store inquiry details, user's emotional state

[1879] Output: Personalized query response

[1880] Through the above processing steps, the system analyzes the user's learning progress and emotional state in real time, providing an optimal learning plan, and enabling personalized and flexible customer service in physical stores.

[1881] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1882] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1883] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1884] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1885] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1886] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1887] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1888] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1889] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1890] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1891] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1892] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1893] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1894] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1895] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1896] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1897] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1898] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1899] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1900] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1901] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1902] The following is further disclosed regarding the above embodiment.

[1903] (Claim 1)

[1904] A means for inputting user information and generating a study plan based on the information;

[1905] a means for conducting a multilingual conversation using a natural language processing engine based on the generated learning plan;

[1906] means for assessing the user's progress and modifying or adjusting the learning plan based on said assessment;

[1907] A means of providing activities for practicing listening, speaking, reading, and writing skills based on the generated lesson plan; and

[1908] A system including:

[1909] (Claim 2)

[1910] 10. The system of claim 1, further comprising: means for registering user information in a database and generating a study plan.

[1911] (Claim 3)

[1912] 10. The system of claim 1, further comprising: means for interacting with a user in real time, the means including a robotic device.

[1913] "Example 1"

[1914] (Claim 1)

[1915] A means for inputting user information and generating a study plan based on the information;

[1916] a means for conducting a multilingual conversation using a natural language processing engine based on the generated learning plan;

[1917] means for assessing the user's progress and modifying or adjusting the study plan based on said assessment;

[1918] A means for providing activities for practicing auditory comprehension, speaking, reading, and writing skills based on the generated lesson plan;

[1919] A system including:

[1920] (Claim 2)

[1921] 10. The system of claim 1, further comprising: means for registering user information in a database and generating a study plan.

[1922] (Claim 3)

[1923] 10. The system of claim 1, further comprising means including an automated device for interacting with the user in real time.

[1924] "Application Example 1"

[1925] (Claim 1)

[1926] A means for inputting user information and generating a study plan based on the information;

[1927] a means for conducting a multilingual conversation using a natural language processing engine based on the generated learning plan;

[1928] means for assessing the user's progress and modifying or adjusting the learning plan based on said assessment;

[1929] A means of providing activities for practicing listening, speaking, reading, and writing skills based on the generated lesson plan; and

[1930] a means of providing multilingual learning using displays, microphones, and speakers within the vehicle;

[1931] A system including:

[1932] (Claim 2)

[1933] 10. The system of claim 1, further comprising: means for registering user information in a database and generating a study plan.

[1934] (Claim 3)

[1935] 10. The system of claim 1, further comprising: means for interacting with a user in real time, the means including a robotic device and an in-vehicle display device.

[1936] "Example 2: Combining Emotion Engines"

[1937] (Claim 1)

[1938] A means for inputting user information and generating a study plan based on the information;

[1939] a means for conducting a multilingual conversation using a natural language processing engine based on the generated learning plan;

[1940] means for assessing the user's progress and modifying or adjusting the learning plan based on said assessment;

[1941] A means to harmoniously train listening, speaking, reading, and writing skills,

[1942] means for analyzing a user's emotional state and dynamically adjusting the learning plan;

[1943] A system including:

[1944] (Claim 2)

[1945] 10. The system of claim 1, further comprising: means for registering user information in a database and generating a study plan.

[1946] (Claim 3)

[1947] 10. The system of claim 1, further comprising: means for interacting with a user in real time, the means including an information processing device.

[1948] "Application example 2 when combining emotion engines"

[1949] (Claim 1)

[1950] A means for inputting user information and generating a study plan based on the information;

[1951] a means for conducting a multilingual conversation using a natural language processing engine based on the generated learning plan;

[1952] means for assessing the user's progress and modifying or adjusting the learning plan based on said assessment;

[1953] A means of providing activities for practicing listening, speaking, reading, and writing skills based on the generated lesson plan; and

[1954] means for analyzing the emotional state of a user and providing an optimal response in real time, the means including an emotion analysis engine;

[1955] a means including an information processing device for handling customer inquiries in a store;

[1956] A system including:

[1957] (Claim 2)

[1958] 10. The system of claim 1, further comprising: means for registering user information in a database and generating a study plan.

[1959] (Claim 3)

[1960] 10. The system of claim 1, further comprising: means for interacting with a user in real time, the means including a robotic device. [Explanation of symbols]

[1961] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for inputting user information and generating a study plan based on the information; a means for conducting a multilingual conversation using a natural language processing engine based on the generated learning plan; means for assessing the user's progress and modifying or adjusting the learning plan based on said assessment; A means of providing activities for practicing listening, speaking, reading, and writing skills based on the generated lesson plan; and A system including:

2. 10. The system of claim 1, further comprising means for registering user information in a database and generating a study plan.

3. 10. The system of claim 1, further comprising: means including a robotic device for interacting with a user in real time.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A