system

The system addresses the limitations of conventional language learning by generating personalized plans, providing real-time voice feedback, and integrating cultural content, resulting in effective language skill improvement and practical application.

JP2026071668APending Publication Date: 2026-04-30SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024181706
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-10-17
Publication Date
2026-04-30

AI Technical Summary

Technical Problem

Conventional language learning systems fail to provide personalized learning plans, real-time pronunciation and grammar feedback, and sufficient cultural integration, limiting the effectiveness of language acquisition and practical application.

Method used

A system that generates individualized learning plans, provides real-time voice analysis and feedback, facilitates collaborative learning, and integrates cultural information and media content to enhance language learning experience.

Benefits of technology

Enables efficient improvement of language skills by offering tailored learning experiences, real-time feedback, and cultural understanding, simulating individual tutoring and enhancing practical language application.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026071668000001_ABST
    Figure 2026071668000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A learning plan generation means that generates a language learning plan based on user input, A voice input receiving means that receives voice input corresponding to a conversation scenario provided in accordance with the learning plan, An analysis means that analyzes the aforementioned voice input in real time and generates feedback regarding pronunciation and grammar, A means for presenting the aforementioned feedback to the user, An adjustment means for adjusting the learning plan based on the user's learning progress, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0004] , , ,

[0005] , , , ,

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In conventional language learning systems, it is difficult to formulate a learning plan according to individual goals and levels of users, and there is a problem that the improvement of realistic conversation ability cannot be efficiently achieved. Also, pronunciation and grammar guidance in real time are not sufficient, and furthermore, insufficient media integration for users to learn while understanding cultural backgrounds is cited as a problem. As a result, an environment that maximally brings out the effect of language learning is not provided.

Means for Solving the Problems

[0005] This invention provides means for generating an appropriate learning plan based on user input, and means for providing real-time voice analysis and feedback. This system flexibly adjusts the learning plan according to the user's learning progress and further enables collaborative language practice with other learners using social learning methods. In addition, by integrating cultural information and media content related to the language the user is learning, it provides a more comprehensive and effective learning experience. This allows users to achieve a learning experience similar to individual tutoring, enabling them to improve their language skills efficiently and effectively.

[0006] A "learning plan generation means" is a device or system that has the function of generating an individually designed learning plan based on the user's language learning goals and level.

[0007] A "voice input receiving means" is a device or system that receives data entered by a user by voice and passes it to the system.

[0008] "Analysis means" refers to a device or system that has the function of analyzing pronunciation and grammar based on received voice input and providing feedback to the user.

[0009] A "presentation means" refers to a device or system for effectively presenting feedback generated based on analysis results to the user.

[0010] "Adjustment means" refers to devices or systems for appropriately updating or modifying existing learning plans according to the user's learning progress.

[0011] A "social learning tool" is a device or system that has the function of supporting multiple users to connect with each other and practice language together.

[0012] A "media integration device" is a device or system that has the function of integrating cultural information and media content related to the user's learning language and providing it to the user. [Brief explanation of the drawing]

[0013] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]

[0014] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0015] First, the terms used in the following description will be explained.

[0016] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0017] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0018] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0019] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0021] [First Embodiment]

[0022] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0023] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0024] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0025] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0026] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0028] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0029] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0030] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0031] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0032] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0033] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0034] This invention is a system that supports users' language learning and consists of a learning plan generation means, a voice input receiving means, an analysis means, a presentation means, an adjustment means, a social learning means, and a media integration means. Each means works in conjunction with the others to enable the provision of instruction tailored to the individual learning needs of the user.

[0035] First, the user inputs their language learning goals and current language level using a device. This information is sent from the device to the server. Based on this, the server generates a personalized learning plan using a learning plan generation mechanism. This learning plan includes necessary content and assignments and is dynamically adjusted according to the user's learning progress.

[0036] Next, the user responds verbally to the conversation scenario provided through the terminal. A voice input receiving device receives this voice input and sends it to the server. The server uses analysis tools, including voice analysis technology, to analyze the user's pronunciation and grammar in real time. The results of this analysis are generated as feedback regarding pronunciation accuracy and grammatical correctness.

[0037] The device presents this feedback to the user using a presentation method. The feedback is provided in text or audio format to aid the user's understanding.

[0038] Furthermore, social learning tools allow users to connect with other learners online and engage in collaborative language practice and discussions. This feature promotes interaction with learners worldwide and enhances social skills.

[0039] Furthermore, through media integration, the server selects and provides cultural background information, news articles, and videos related to the language the user is learning. This feature helps users learn the language in a practical context and contributes to deepening their cultural understanding.

[0040] As a concrete example, let's say User B is learning English with the goal of traveling. The terminal presents a conversation scenario at a check-in counter, and the user inputs, "Can I have my boarding pass, please?" The server analyzes the speech and provides appropriate response examples along with feedback on accurate pronunciation. In addition, by visually presenting videos about British culture and airport etiquette, the user can simultaneously acquire knowledge beyond pronunciation and grammar.

[0041] In this way, this system comprehensively supports language learning and effectively assists users in achieving their goals.

[0042] The following describes the processing flow.

[0043] Step 1:

[0044] Users launch a language learning application on their device and set individual language learning goals, such as "travel" or "business." They also select their current language level and take a self-assessment test to input their detailed skill level.

[0045] Step 2:

[0046] The terminal sends user input information to the server. This information includes learning objectives and language level assessment results.

[0047] Step 3:

[0048] Based on the received information, the server generates a user-specific learning plan using a learning plan generation mechanism. The generated plan, which includes specific conversation scenarios and related tasks, is sent to the user's terminal.

[0049] Step 4:

[0050] The user follows the instructions for pronunciation practice and makes voice input according to the conversation scenario displayed on the device. When the user reads a sample sentence aloud, the device sends the audio to the server.

[0051] Step 5:

[0052] The server receives voice input and uses analysis tools to evaluate pronunciation, grammar, and fluency. Speech recognition technology and AI-based natural language processing are used for the analysis, and evaluation results are generated.

[0053] Step 6:

[0054] The server generates and constructs feedback, which is then sent to the terminal. This feedback includes corrections to pronunciation and suggestions for improvement, helping the user learn.

[0055] Step 7:

[0056] The device displays the feedback it receives to the user. The feedback can be presented as text or audio, allowing the user to track their progress and move on to the next step.

[0057] Step 8:

[0058] Users can connect online with other learners through social learning tools. This allows users to practice conversation, share cultural information, and advance their learning further.

[0059] Step 9:

[0060] The server collects cultural information and the latest news related to the language the user is learning and transmits it to the terminal through a media integration system. This allows users to learn not only the language but also the culture behind it.

[0061] Step 10:

[0062] The user's learning progress is tracked over a certain period, and the server dynamically adjusts the learning plan based on this data, suggesting new challenges and learning materials. This improves the effectiveness and continuity of learning.

[0063] (Example 1)

[0064] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0065] Modern language education demands personalized learning plans and real-time analysis of speech input to provide feedback. However, existing systems are not sufficiently individualized, particularly in improving practical language application skills and cultural understanding. Furthermore, collaborative learning with other learners and the integration of relevant media have not been adequately achieved to enhance learning effectiveness. Therefore, the challenge lies in enabling users to effectively acquire language skills and apply them in real-world situations.

[0066] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0067] In this invention, the server includes a plan generation means, an audio data receiving means, and a data analysis means. This enables the generation of a personalized education plan that is dynamically adjusted according to the user's progress, real-time analysis of voice input, and provision of specific pronunciation and grammar correction instructions. Furthermore, data is sent and received effectively through secure communication means. In addition, collaborative learning means allow users to work with other educators and enhance their language skills in a practical setting. Moreover, media integration means allow users to gain a deeper understanding through relevant cultural background and media information.

[0068] "Plan generation means" refers to a function for creating personalized educational plans based on user input information.

[0069] "Audio data receiving means" refers to a function that receives audio information transmitted by the user on a terminal and sends it to a server.

[0070] "Data analysis means" refers to a function that instantly analyzes received audio information and evaluates the user's pronunciation and grammar.

[0071] "Information presentation means" refers to functions that present feedback generated by the server to the user to aid in understanding.

[0072] "Plan adjustment mechanism" refers to a function that dynamically modifies the educational plan according to the user's learning progress.

[0073] "Communication method" refers to the function of using protocols to securely send and receive data between a server and a terminal.

[0074] "Feedback generation means" refers to a function that evaluates the accuracy of pronunciation and grammar to the user based on the analyzed results and provides instructions.

[0075] "Collaborative learning tools" refer to features that allow users to collaborate with other educators and practice language together.

[0076] "Media integration means" refers to functions that integrate and provide cultural information and media content related to the language a user is learning.

[0077] This invention is a system for supporting users' language education, and it functions through the coordinated action of multiple components. Specific embodiments are described below.

[0078] First, the user enters their target language and current language level using a terminal. A dedicated application software runs on the terminal, accepting input through a user interface. The entered data is securely transmitted from the terminal to the server. This transmission utilizes encryption technologies such as the SSL / TLS protocol.

[0079] The server generates personalized learning plans using a plan generation mechanism based on the received user data. This generation process utilizes algorithms, for example, implemented in Python. The generated plans are dynamically adjusted based on the user's language learning progress. The learning plans include scenarios specifically tailored to different skills such as listening, speaking, and reading.

[0080] Next, the user uses their device to provide voice responses to a specified scenario. This voice is converted into a digital signal by the device and transmitted to the server via an audio data receiving device. The server instantly analyzes the voice data using an analysis device and evaluates the pronunciation and grammar. Specific voice analysis utilizes speech recognition APIs and natural language processing (NLP) technologies.

[0081] Based on the analysis results, the server uses a feedback generation mechanism to create a response for the user. This response includes accuracy of pronunciation and grammatical corrections. The terminal then uses an information presentation mechanism to present this feedback to the user in an easy-to-understand manner. The feedback is provided in both text and audio formats.

[0082] Furthermore, through collaborative learning methods, users can connect with other learners in real time and practice language together. Additionally, media integration allows the server to extract cultural background and media information and provide it to users through their terminals. This enables users to deepen not only their language acquisition but also their understanding of culture.

[0083] For example, if a user sets "Useful English Phrases for Travel" as their learning goal, the device will present a conversation scenario at an airport check-in counter. The user will say, "Can I have my boarding pass, please?" The server will analyze this audio and provide appropriate feedback. Furthermore, as relevant cultural background, a video about airport etiquette will be played on the device.

[0084] Examples of prompt messages are as follows:

[0085] "To help users practice everyday conversation, please set up situations and provide feedback on pronunciation and grammar."

[0086] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0087] Step 1:

[0088] The user launches the application on the terminal and enters the target language and current language level. During this process, text data is entered through the user interface, and the terminal converts it into a digital format. The entered information is securely transmitted to the server. The output is the user information data sent to the server.

[0089] Step 2:

[0090] The server analyzes the received user information and generates a personalized educational plan using a plan generation mechanism. An algorithm implemented in Python specifies appropriate tasks and content based on the user's goals. The input is digital data of the user information, and the output is the generated educational plan. This plan is stored on the server and sent to the terminal.

[0091] Step 3:

[0092] The user responds verbally to conversation scenarios provided via the terminal. The terminal uses a speech recognition device to convert the speech into a digital signal and sends this audio data to the server. The input is the user's voice data, and the output is the digital audio signal transferred to the server.

[0093] Step 4:

[0094] The server processes the received audio data using analysis tools. Specifically, it converts the audio to text using a speech recognition API and evaluates the grammar and pronunciation using natural language processing techniques. It generates feedback based on the analysis results. The input is a digital audio signal, and the output is the analysis results and the feedback data based on them.

[0095] Step 5:

[0096] The terminal presents the user with feedback received from the server. This feedback includes accuracy of pronunciation, grammatical corrections, and areas for improvement, and is presented to the user visually and audibly. The input is the feedback data, and the output is the feedback information presented to the user.

[0097] Step 6:

[0098] Users connect with other learners using collaborative learning tools to practice and discuss together. The device enables real-time interaction via an online platform. Input is communication data from other learners, and output is user interaction.

[0099] Step 7:

[0100] The server uses media integration to select cultural content relevant to the user's learning and provides it to the user through the terminal. This allows the user to learn language and related culture simultaneously. The input is data of the cultural content, and the output is media information presented to the user.

[0101] (Application Example 1)

[0102] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0103] Conventional language learning systems have struggled to provide flexible learning methods tailored to individual learning needs. Furthermore, they lacked sufficient mechanisms for learners to check and correct their pronunciation and grammatical errors in real time. Additionally, methods for learners to deepen their understanding of cultural context relevant to actual usage environments were limited. Therefore, the present invention aims to provide an individualized learning experience and improve users' multifaceted language abilities.

[0104] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0105] In this invention, the server includes planning means for generating a language learning plan based on user input, speech acquisition means for analyzing speech related to vocabulary training in real time, and material provision means for providing materials to the user through a visual medium. This allows the user to check and correct their pronunciation and grammar on the spot through always personalized, real-time feedback. It also facilitates understanding of cultural background through a visual medium.

[0106] A "planning mechanism" is a system for generating personalized language learning plans based on user input information.

[0107] A "voice acquisition means" is a mechanism for receiving and analyzing the voice emitted by the user in real time.

[0108] The "analysis means" is a mechanism for detecting pronunciation and grammatical errors based on received audio data and generating feedback.

[0109] A "display means" is a mechanism for presenting feedback generated by analysis to the user in text or audio format.

[0110] The "adjustment mechanism" is a system for updating and optimizing the generated learning plan in a timely manner according to the user's learning progress.

[0111] A "means of providing information" refers to a mechanism that provides users with materials including visual information and cultural background to promote a deeper understanding.

[0112] A "media integration tool" is a mechanism for integrating cultural information related to language learning and providing it to learners.

[0113] A "collaborative learning tool" is a mechanism that allows users to work together with other learners to train and practice language collaboratively.

[0114] The system for realizing this invention integrates a variety of functions to effectively support the user's language learning. Next, the operation of this system will be described.

[0115] Users first input their learning goals and current language level using a device such as a smartphone or tablet. The device then sends this information to the server. The server, as a planning tool, generates a personalized learning plan based on this user information. This plan can be flexibly modified according to the user's progress.

[0116] When a user practices speaking, the system receives the spoken audio in real time through the audio acquisition mechanism. The server processes this audio data using an analysis mechanism and generates feedback on pronunciation and grammatical errors. The display mechanism presents the feedback to the user to aid understanding.

[0117] Furthermore, the materials provided offer users visual information, such as cultural background and related video materials. This allows users to learn not only about language but also about the cultural aspects of the language in depth.

[0118] This system is implemented as a smartphone application using React Native, with AWS Lambda running on a cloud server. AWS Transcribe and PyDub are used for speech analysis. Feedback is synthesized using AWS Polly.

[0119] As a concrete example, a user might watch a scene from a movie and practice pronouncing the lines from it. In this case, the server analyzes the pronunciation and provides appropriate feedback in real time. Furthermore, it can provide the cultural background and related information of the scene as video material.

[0120] An example of a prompt sentence would be, "I want to watch a scene from a movie and try to pronounce it myself to get accurate feedback and cultural understanding."

[0121] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0122] Step 1:

[0123] The user inputs their learning goals and current language level from their device. The device sends this data to the server. The server receives the input data and uses a planning tool to generate a personalized learning plan tailored to the user. This plan includes the content and resources necessary for the user to achieve their immediate learning goals.

[0124] Step 2:

[0125] The user uses a device to perform voice input as part of language learning. The device captures voice data through its microphone and sends it to the server via a voice acquisition device. The server converts the received voice into text data using AWS Transcribe and performs detailed voice analysis using PyDub. Here, the waveform and phonemes of the voice data are analyzed to generate feedback on pronunciation and grammar.

[0126] Step 3:

[0127] The server sends the feedback obtained from the analysis to the terminal via a display device. The terminal displays this feedback on the screen in text format, or, if necessary, communicates it to the user in audio format using AWS Polly. The user can use this information to gain a better understanding of where improvements are needed.

[0128] Step 4:

[0129] The server uses a resource delivery system to select visual literature and videos relevant to the user's language learning and transmits this information to the terminal. The terminal displays this information to the user, facilitating an understanding of the language's cultural background and related topics. This allows the user to acquire a broad range of knowledge beyond mere language skills.

[0130] Step 5:

[0131] As users progress through their learning plan, they update their learning status based on new challenges and what they have learned. The server uses this feedback to periodically revise the learning plan using adjustment mechanisms, providing the user with the most optimal route. This enables flexible learning tailored to their progress.

[0132] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0133] This invention provides a system for more effective language learning, generating individualized learning plans for users and offering an interactive experience. In addition to learning plan generation means, voice input receiving means, analysis means, presentation means, and adjustment means, this system incorporates an emotion engine. This allows for the provision of a customized learning experience that takes user emotions into consideration.

[0134] First, the user accesses the application via their device and inputs their learning goals and current skill level. This generates a personalized learning plan. The device sends this information to the server, which then develops the learning plan.

[0135] The user initiates an interactive conversation by using voice input through their device. The emotion engine then analyzes the user's tone and word choice to recognize their emotions. This emotional information is used to further personalize the learning experience.

[0136] For example, if the server detects that the user is feeling stressed, it can provide relaxing scenarios or encouraging feedback to alleviate this. Conversely, if the user is enjoying themselves, the server can increase the difficulty level and provide more challenging tasks.

[0137] The generated learning feedback is presented to the user in audio or text format through various presentation methods. Furthermore, to facilitate interaction with other learners, the user's emotional state is used in an anonymized form, and the next steps are planned through social learning tools.

[0138] As a concrete example, when user C is learning a language, the emotion engine detects from the user's voice that they are "satisfied." The server uses this emotion information to provide user C with the next level of advanced conversation scenarios, helping them improve their skills while maintaining their motivation.

[0139] Thus, the present invention allows users to receive an optimal learning experience tailored to their mood and state at any given time, dramatically improving the efficiency and satisfaction of language acquisition.

[0140] The following describes the processing flow.

[0141] Step 1:

[0142] The user starts a session on their device and enters their learning goal (e.g., conversational English for travel) and current language level. This information is necessary to create a customized learning plan to meet the user's individual needs.

[0143] Step 2:

[0144] The terminal sends user input information to the server. The server uses a learning plan generation mechanism to generate a specific learning plan based on the user's goals and level. This plan includes learning content and conversation scenarios.

[0145] Step 3:

[0146] The user provides voice input in response to a conversation scenario displayed on their device. The voice input receiving device captures this audio and sends it to the server.

[0147] Step 4:

[0148] The server processes the received audio using analysis tools and an emotion engine. During this process, the user's pronunciation and grammar are evaluated, while the emotion engine recognizes the user's emotional state (e.g., tension, joy, concentration) from the audio.

[0149] Step 5:

[0150] Based on the analysis results and recognized emotions, the server generates feedback. This feedback includes not only comments on pronunciation and grammar, but also encouragement and advice tailored to the user's emotions. For example, if tension is detected, simple exercises to relax may be recommended.

[0151] Step 6:

[0152] Through a presentation mechanism, the terminal presents feedback from the server to the user. The feedback is displayed as text or audio, which the user reviews and understands.

[0153] Step 7:

[0154] Based on the feedback, users can choose to continue with the next conversation practice or connect with other learners to utilize social learning methods. In this process, the user's emotional state is shared anonymously and used to facilitate interaction with other learners.

[0155] Step 8:

[0156] The server monitors the user's learning progress, adaptively adjusts the learning plan, and suggests new content and scenarios. This helps maintain the user's continuous learning and motivation.

[0157] This process enables the system to provide an optimal learning experience tailored to each user's individual needs, maximizing the effectiveness of language acquisition.

[0158] (Example 2)

[0159] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0160] In today's learning environment, there is a challenge in providing appropriate learning experiences tailored to the individual state and emotions of each learner. In particular, traditional learning methods do not adequately provide individualized feedback and adjustments based on the user's emotions and progress, and there is a need for efficient methods to maximize learning effectiveness.

[0161] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0162] In this invention, the server includes a plan generation means for generating a learning plan based on user input, an input receiving means for receiving voice input, an analysis means for analyzing the voice input and generating feedback, and an emotion analysis means for analyzing the user's emotional state and adjusting the learning experience. This makes it possible to provide the most appropriate learning experience for each individual user in real time.

[0163] The "plan generation means" is a function that generates an individualized learning plan based on the user's input information.

[0164] The "input receiving means" refers to a function that receives voice input from the user.

[0165] The "analysis means" refers to a function that analyzes received audio data and generates feedback related to pronunciation and grammar.

[0166] A "presentation method" is a function that provides the generated feedback to the user visually or audibly.

[0167] The "adjustment mechanism" is a function that dynamically adjusts the existing learning plan based on the user's learning progress.

[0168] "Emotional analysis tools" are functions that analyze the tone of the user's voice and word choices to recognize the user's emotional state.

[0169] This invention is a language learning system designed to enhance the user's individual learning experience. This system can generate dynamic learning plans tailored to the user's individual needs and adjust the learning experience based on emotions.

[0170] The user first accesses a language learning application through their device and inputs their learning goals and current proficiency level. The device sends this information to a server. The server uses a generative AI model to generate a personalized learning plan based on the provided data. This plan is then executed interactively via voice input.

[0171] The voice input from the user via the device is sent to the server by an input receiving device. The server analyzes the user's emotional state using an emotion analysis device and provides feedback and learning adjustments accordingly. A generative AI model provides optimal learning content based on the user's current learning status and emotions. Specifically, this includes providing a challenging task when the user is judged to be relaxed, or providing encouraging messages when the user is stressed.

[0172] The presentation methods allow users to visually or audibly confirm feedback and the next learning steps. This enables learners to progress at their own pace. Furthermore, by incorporating social learning methods that facilitate interaction with other learners, it is possible to further enrich the diversity of learning.

[0173] For example, if a user inputs "Teach me a new word" via voice, the server can select a word appropriate to its difficulty level and present it along with an example sentence. Another example of a prompt is, "Please use the emotion engine to provide appropriate feedback so that the user can learn in a relaxed state." In this way, it is possible to provide an optimal learning experience tailored to the user's emotions and state.

[0174] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0175] Step 1:

[0176] Users access the learning platform via their device and log in. After logging in, users enter their learning goals (e.g., mastering everyday conversation) and current proficiency level (e.g., beginner, intermediate, advanced). The entered data is sent from the device to the server and used as input to create a personalized learning plan.

[0177] Step 2:

[0178] The server uses a generative AI model to generate a personalized learning plan based on the user's learning goals and ability level. The generative AI model considers the user's past learning history and preference patterns to select the most suitable learning materials and assignments. This generates a learning plan that matches the user's needs and sends it to the device.

[0179] Step 3:

[0180] The user uses a device to input voice data and initiate an interactive learning session. The device receives the user's voice data in real time and sends it to the server. The voice data mainly consists of phrases and questions used during the learning process.

[0181] Step 4:

[0182] The server analyzes the audio data received from the terminal and generates feedback on pronunciation and grammar. The audio analysis also includes sentiment analysis, determining the user's emotions from their tone of voice and word choice. The generated feedback is adjusted according to the user's emotional state.

[0183] Step 5:

[0184] Based on the results of the emotion analysis, the server dynamically adjusts the user's learning experience. For example, if the server detects tension, it provides calming feedback and easier tasks; if it determines that the user is enjoying themselves, it offers more challenging tasks. The adjusted feedback is then sent from the server to the user's device.

[0185] Step 6:

[0186] The device presents the user with feedback and adjusted learning materials sent from the server. This is often done through text display or audio output. The user then uses this feedback to advance their learning.

[0187] Step 7:

[0188] If users wish to engage in further interaction or learning, they can repeat the same process. Furthermore, social learning features allow users to deepen their interactions with other learners, providing a more diverse learning experience.

[0189] (Application Example 2)

[0190] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0191] One challenge for users is that it's difficult to receive product recommendations that match their emotions and current mood when selecting products. This leads to information overload in the product selection process, causing users to struggle to make decisions.

[0192] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0193] In this invention, the server includes an emotion analysis device that analyzes the user's emotions and adjusts suggestions, a collaborative work device, and an information integration device. This makes it possible to provide optimal product information based on the user's emotions.

[0194] A "plan generation device" is a device that automatically creates an optimal learning and recommendation plan based on the user's input information.

[0195] An "input receiving device" is a device for receiving voice and text data provided by the user.

[0196] An "analysis device" is a device that analyzes received data and generates necessary information and feedback.

[0197] A "presentation device" is a device used to present feedback and information generated by an analysis device to the user.

[0198] A "adjustment device" is a device that optimizes and adjusts the plan according to the user's progress and emotions.

[0199] An "emotion analysis device" is a device that analyzes the user's current emotions from their voice or text and adjusts the operation of the entire system based on the results.

[0200] A "collaborative work device" is a device that provides an environment for performing tasks and learning in cooperation with other users.

[0201] An "information integration device" is a device that integrates various types of information related to the user and provides them in an appropriate format.

[0202] To implement this invention, an advanced interactive system is constructed to provide information and learning content tailored to the user's emotions. The system consists of a plan generation device, an input receiving device, an analysis device, a presentation device, an adjustment device, an emotion analysis device, a collaborative work device, and an information integration device.

[0203] The plan generation device generates learning plans and product recommendation plans that match the user's requests based on data entered by the user using a terminal. For example, suppose a user enters "I want to try a new language learning method" through a smartphone application. This input is sent to the server, received by the plan generation device, and provides a plan that suits the user's needs.

[0204] The input receiving device receives user voice and text input. In this system, for example, voice input is converted to text using the Google® Speech-to-Text API. This converted text is then moved to the next parsing stage.

[0205] The analysis device analyzes the received data to determine pronunciation, grammar, and product review recommendation information.

[0206] Alternatively, it generates feedback. IBM Watson® NLU, among others, is used for analysis.

[0207] The display device provides the user with visual or audible feedback on the results. This feedback is tailored to the user's preferences, such as presenting entertainment information for a visit to a restaurant with a strong Latin American theme.

[0208] The adjustment device fine-tunes the plan according to the user's progress and emotions. For example, if the user is feeling anxious, it will provide materials with a lower difficulty level.

[0209] The emotion analysis device reads the user's emotions from their voice and facial expressions. This is extremely useful in determining the system's subsequent behavior. The information presented is adjusted accordingly depending on whether the voice tone is tense or relaxed.

[0210] Collaborative work systems can link data when multiple users are viewing it simultaneously. For example, they can share opinions when a group of users are selecting a product together.

[0211] Information integration devices have the function of integrating information relevant to the user, thereby providing a highly satisfying experience. They retrieve the most relevant information from the database and construct a valuable experience for the user.

[0212] This design allows users to intuitively obtain the information they desire. An example of a prompt would be, "Show me reviews that will put my mind at ease." This prompt provides information that is sensitive to the user's emotional state.

[0213] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0214] Step 1:

[0215] The device receives voice input from the user. The user begins speaking into the device, and this is captured as audio data by the device's microphone. This audio data becomes the entry point for the system.

[0216] Step 2:

[0217] The device converts the captured audio data into text data using the Google Speech-to-Text API. The input here is audio data, and the output is text data. This converted text then proceeds to the next analysis phase.

[0218] Step 3:

[0219] The server analyzes text data using IBM Watson NLU to extract user emotions and intentions. The input is the text data generated in step 2, and the output is data on the user's emotions and specific intentions. This enables suggestions tailored to the user's emotions.

[0220] Step 4:

[0221] Based on the analyzed emotional information, the server's planning generator produces optimal recommendations and learning content for the user. The input here is emotional and intent data, and the output is a specific learning plan or product recommendation plan.

[0222] Step 5:

[0223] The server provides feedback to the user through a presentation device, which presents the generated plan. The information presented is communicated to the user in visual or audio format. The input is the plan generated in step 4, and the output is the feedback displayed or played back to the user.

[0224] Step 6:

[0225] Based on user feedback and usage history, the adjustment system fine-tunes the plan. The server uses this information to make suggestions more user-friendly for subsequent interactions. This cycle continuously improves the user experience.

[0226] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0227] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0228] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0229] [Second Embodiment]

[0230] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0231] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0232] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0233] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0234] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0235] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0236] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0237] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0238] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0239] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0240] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0241] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0242] This invention is a system that supports users' language learning and consists of a learning plan generation means, a voice input receiving means, an analysis means, a presentation means, an adjustment means, a social learning means, and a media integration means. Each means works in conjunction with the others to enable the provision of instruction tailored to the individual learning needs of the user.

[0243] First, the user inputs their language learning goals and current language level using a device. This information is sent from the device to the server. Based on this, the server generates a personalized learning plan using a learning plan generation mechanism. This learning plan includes necessary content and assignments and is dynamically adjusted according to the user's learning progress.

[0244] Next, the user responds verbally to the conversation scenario provided through the terminal. A voice input receiving device receives this voice input and sends it to the server. The server uses analysis tools, including voice analysis technology, to analyze the user's pronunciation and grammar in real time. The results of this analysis are generated as feedback regarding pronunciation accuracy and grammatical correctness.

[0245] The device presents this feedback to the user using a presentation method. The feedback is provided in text or audio format to aid the user's understanding.

[0246] Furthermore, social learning tools allow users to connect with other learners online and engage in collaborative language practice and discussions. This feature promotes interaction with learners worldwide and enhances social skills.

[0247] Furthermore, through media integration, the server selects and provides cultural background information, news articles, and videos related to the language the user is learning. This feature helps users learn the language in a practical context and contributes to deepening their cultural understanding.

[0248] As a concrete example, let's say User B is learning English with the goal of traveling. The terminal presents a conversation scenario at a check-in counter, and the user inputs, "Can I have my boarding pass, please?" The server analyzes the speech and provides appropriate response examples along with feedback on accurate pronunciation. In addition, by visually presenting videos about British culture and airport etiquette, the user can simultaneously acquire knowledge beyond pronunciation and grammar.

[0249] In this way, this system comprehensively supports language learning and effectively assists users in achieving their goals.

[0250] The following describes the processing flow.

[0251] Step 1:

[0252] Users launch a language learning application on their device and set individual language learning goals, such as "travel" or "business." They also select their current language level and take a self-assessment test to input their detailed skill level.

[0253] Step 2:

[0254] The terminal sends user input information to the server. This information includes learning objectives and language level assessment results.

[0255] Step 3:

[0256] Based on the received information, the server generates a user-specific learning plan using a learning plan generation mechanism. The generated plan, which includes specific conversation scenarios and related tasks, is sent to the user's terminal.

[0257] Step 4:

[0258] The user follows the instructions for pronunciation practice and makes voice input according to the conversation scenario displayed on the device. When the user reads a sample sentence aloud, the device sends the audio to the server.

[0259] Step 5:

[0260] The server receives voice input and uses analysis tools to evaluate pronunciation, grammar, and fluency. Speech recognition technology and AI-based natural language processing are used for the analysis, and evaluation results are generated.

[0261] Step 6:

[0262] The server generates and constructs feedback, which is then sent to the terminal. This feedback includes corrections to pronunciation and suggestions for improvement, helping the user learn.

[0263] Step 7:

[0264] The device displays the feedback it receives to the user. The feedback can be presented as text or audio, allowing the user to track their progress and move on to the next step.

[0265] Step 8:

[0266] Users can connect online with other learners through social learning tools. This allows users to practice conversation, share cultural information, and advance their learning further.

[0267] Step 9:

[0268] The server collects cultural information and the latest news related to the language the user is learning and transmits it to the terminal through a media integration system. This allows users to learn not only the language but also the culture behind it.

[0269] Step 10:

[0270] The user's learning progress is tracked over a certain period, and the server dynamically adjusts the learning plan based on this data, suggesting new challenges and learning materials. This improves the effectiveness and continuity of learning.

[0271] (Example 1)

[0272] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0273] Modern language education demands personalized learning plans and real-time analysis of speech input to provide feedback. However, existing systems are not sufficiently individualized, particularly in improving practical language application skills and cultural understanding. Furthermore, collaborative learning with other learners and the integration of relevant media have not been adequately achieved to enhance learning effectiveness. Therefore, the challenge lies in enabling users to effectively acquire language skills and apply them in real-world situations.

[0274] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0275] In this invention, the server includes a plan generation means, an audio data receiving means, and a data analysis means. This enables the generation of a personalized education plan that is dynamically adjusted according to the user's progress, real-time analysis of voice input, and provision of specific pronunciation and grammar correction instructions. Furthermore, data is sent and received effectively through secure communication means. In addition, collaborative learning means allow users to work with other educators and enhance their language skills in a practical setting. Moreover, media integration means allow users to gain a deeper understanding through relevant cultural background and media information.

[0276] "Plan generation means" refers to a function for creating personalized educational plans based on user input information.

[0277] "Audio data receiving means" refers to a function that receives audio information transmitted by the user on a terminal and sends it to a server.

[0278] "Data analysis means" refers to a function that instantly analyzes received audio information and evaluates the user's pronunciation and grammar.

[0279] "Information presentation means" refers to functions that present feedback generated by the server to the user to aid in understanding.

[0280] The "Plan Adjustment Means" refers to a function for dynamically modifying an educational plan according to the learning progress of the user.

[0281] The "Communication Means" refers to a function that uses a protocol for securely transmitting and receiving data between the server and the terminal.

[0282] The "Feedback Generation Means" refers to a function that evaluates the accuracy of pronunciation and grammar for the user based on the analyzed results and provides instructions.

[0283] The "Collaborative Learning Means" refers to a function for the user to cooperate with other educators and jointly conduct language practice.

[0284] The "Media Integration Means" refers to a function for integrating and providing cultural information and media content related to the language the user is learning.

[0285] The present invention is a system for assisting the language education of users, in which a plurality of components cooperate to function. Hereinafter, specific embodiments thereof will be described.

[0286] First, the user inputs the target language and the current language level using the terminal. Dedicated application software operates on the terminal, and this software receives the input via the user interface. The input data is securely transmitted from the terminal to the server. For this transmission, encryption technologies such as the SSL / TLS protocol are used.

[0287] The server uses the plan generation means based on the received user data to generate an individualized educational plan. In this generation process, for example, an algorithm implemented in Python is used. The generated plan is dynamically adjusted based on the progress of the user's language learning. The educational plan includes scenarios specialized for different skills such as listening, speaking, and reading.

[0288] Next, the user uses their device to provide voice responses to a specified scenario. This voice is converted into a digital signal by the device and transmitted to the server via an audio data receiving device. The server instantly analyzes the voice data using an analysis device and evaluates the pronunciation and grammar. Specific voice analysis utilizes speech recognition APIs and natural language processing (NLP) technologies.

[0289] Based on the analysis results, the server uses a feedback generation mechanism to create a response for the user. This response includes accuracy of pronunciation and grammatical corrections. The terminal then uses an information presentation mechanism to present this feedback to the user in an easy-to-understand manner. The feedback is provided in both text and audio formats.

[0290] Furthermore, through collaborative learning methods, users can connect with other learners in real time and practice language together. Additionally, media integration allows the server to extract cultural background and media information and provide it to users through their terminals. This enables users to deepen not only their language acquisition but also their understanding of culture.

[0291] For example, if a user sets "Useful English Phrases for Travel" as their learning goal, the device will present a conversation scenario at an airport check-in counter. The user will say, "Can I have my boarding pass, please?" The server will analyze this audio and provide appropriate feedback. Furthermore, as relevant cultural background, a video about airport etiquette will be played on the device.

[0292] Examples of prompt messages are as follows:

[0293] "To help users practice everyday conversation, please set up situations and provide feedback on pronunciation and grammar."

[0294] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0295] Step 1:

[0296] The user launches the application on the terminal and enters the target language and current language level. During this process, text data is entered through the user interface, and the terminal converts it into a digital format. The entered information is securely transmitted to the server. The output is the user information data sent to the server.

[0297] Step 2:

[0298] The server analyzes the received user information and generates a personalized educational plan using a plan generation mechanism. An algorithm implemented in Python specifies appropriate tasks and content based on the user's goals. The input is digital data of the user information, and the output is the generated educational plan. This plan is stored on the server and sent to the terminal.

[0299] Step 3:

[0300] The user responds verbally to conversation scenarios provided via the terminal. The terminal uses a speech recognition device to convert the speech into a digital signal and sends this audio data to the server. The input is the user's voice data, and the output is the digital audio signal transferred to the server.

[0301] Step 4:

[0302] The server processes the received audio data using analysis tools. Specifically, it converts the audio to text using a speech recognition API and evaluates the grammar and pronunciation using natural language processing techniques. It generates feedback based on the analysis results. The input is a digital audio signal, and the output is the analysis results and the feedback data based on them.

[0303] Step 5:

[0304] The terminal presents the feedback received from the server to the user. The feedback includes pronunciation accuracy, grammar corrections, and areas for improvement, and is presented to the user both visually and auditorily. The input is feedback data, and the output is the feedback information presented to the user.

[0305] Step 6:

[0306] The user connects with other learners using collaborative learning means and conducts joint practice and discussions. The terminal enables real-time interaction via an online platform. The input is communication data from other learners, and the output is the interaction between users.

[0307] Step 7:

[0308] The server selects cultural content related to the user's learning using media integration means and provides this to the user via the terminal. As a result, the user can learn the language and related culture simultaneously. The input is data on cultural content, and the output is the media information presented to the user.

[0309] (Application Example 1)

[0310] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".

[0311] In conventional language learning systems, it has been difficult to provide flexible learning methods according to individual learning needs. Also, the mechanism for learners to confirm and correct their pronunciation and grammar mistakes in real time has been insufficient. Furthermore, the methods for learners to deepen their understanding of cultural backgrounds in line with the actual usage environment have been limited. Therefore, the present invention aims to provide an individualized learning experience and improve the user's comprehensive language ability.

[0312] [[ID=​

[0313] In this invention, the server includes planning means for generating a language learning plan based on user input, speech acquisition means for analyzing speech related to vocabulary training in real time, and material provision means for providing materials to the user through a visual medium. This allows the user to check and correct their pronunciation and grammar on the spot through always personalized, real-time feedback. It also facilitates understanding of cultural background through a visual medium.

[0314] A "planning mechanism" is a system for generating personalized language learning plans based on user input information.

[0315] A "voice acquisition means" is a mechanism for receiving and analyzing the voice emitted by the user in real time.

[0316] The "analysis means" is a mechanism for detecting pronunciation and grammatical errors based on received audio data and generating feedback.

[0317] A "display means" is a mechanism for presenting feedback generated by analysis to the user in text or audio format.

[0318] The "adjustment mechanism" is a system for updating and optimizing the generated learning plan in a timely manner according to the user's learning progress.

[0319] A "means of providing information" refers to a mechanism that provides users with materials including visual information and cultural background to promote a deeper understanding.

[0320] A "media integration tool" is a mechanism for integrating cultural information related to language learning and providing it to learners.

[0321] A "collaborative learning tool" is a mechanism that allows users to work together with other learners to train and practice language collaboratively.

[0322] The system for realizing this invention integrates a variety of functions to effectively support the user's language learning. Next, the operation of this system will be described.

[0323] Users first input their learning goals and current language level using a device such as a smartphone or tablet. The device then sends this information to the server. The server, as a planning tool, generates a personalized learning plan based on this user information. This plan can be flexibly modified according to the user's progress.

[0324] When a user practices speaking, the system receives the spoken audio in real time through the audio acquisition mechanism. The server processes this audio data using an analysis mechanism and generates feedback on pronunciation and grammatical errors. The display mechanism presents the feedback to the user to aid understanding.

[0325] Furthermore, the materials provided offer users visual information, such as cultural background and related video materials. This allows users to learn not only about language but also about the cultural aspects of the language in depth.

[0326] This system is implemented as a smartphone application using React Native, with AWS Lambda running on the cloud server. AWS Transcribe and PyDub are used for speech analysis. Feedback is synthesized using AWS Polly.

[0327] As a concrete example, a user might watch a scene from a movie and practice pronouncing the lines from it. In this case, the server analyzes the pronunciation and provides appropriate feedback in real time. Furthermore, it can provide the cultural background and related information of the scene as video material.

[0328] An example of a prompt sentence would be, "I want to watch a scene from a movie and try to pronounce it myself to get accurate feedback and cultural understanding."

[0329] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0330] Step 1:

[0331] The user inputs their learning goals and current language level from their device. The device sends this data to the server. The server receives the input data and uses a planning tool to generate a personalized learning plan tailored to the user. This plan includes the content and resources necessary for the user to achieve their immediate learning goals.

[0332] Step 2:

[0333] The user uses a device to perform voice input as part of language learning. The device captures voice data through its microphone and sends it to the server via a voice acquisition device. The server converts the received voice into text data using AWS Transcribe and performs detailed voice analysis using PyDub. Here, the waveform and phonemes of the voice data are analyzed to generate feedback on pronunciation and grammar.

[0334] Step 3:

[0335] The server sends the feedback obtained from the analysis to the terminal via a display device. The terminal displays this feedback on the screen in text format, or, if necessary, communicates it to the user in audio format using AWS Polly. The user can use this information to gain a better understanding of where improvements are needed.

[0336] Step 4:

[0337] The server uses a resource delivery system to select visual literature and videos relevant to the user's language learning and transmits this information to the terminal. The terminal displays this information to the user, facilitating an understanding of the language's cultural background and related topics. This allows the user to acquire a broad range of knowledge beyond mere language skills.

[0338] Step 5:

[0339] As users progress through their learning plan, they update their learning status based on new challenges and what they have learned. The server uses this feedback to periodically revise the learning plan using adjustment mechanisms, providing the user with the most optimal route. This enables flexible learning tailored to their progress.

[0340] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0341] This invention provides a system for more effective language learning, generating individualized learning plans for users and offering an interactive experience. In addition to learning plan generation means, voice input receiving means, analysis means, presentation means, and adjustment means, this system incorporates an emotion engine. This allows for the provision of a customized learning experience that takes user emotions into consideration.

[0342] First, the user accesses the application via their device and inputs their learning goals and current skill level. This generates a personalized learning plan. The device sends this information to the server, which then develops the learning plan.

[0343] The user initiates an interactive conversation by using voice input through their device. The emotion engine then analyzes the user's tone and word choice to recognize their emotions. This emotional information is used to further personalize the learning experience.

[0344] For example, if the server detects that the user is feeling stressed, it can provide relaxing scenarios or encouraging feedback to alleviate this. Conversely, if the user is enjoying themselves, the server can increase the difficulty level and provide more challenging tasks.

[0345] The generated learning feedback is presented to the user in audio or text format through various presentation methods. Furthermore, to facilitate interaction with other learners, the user's emotional state is used in an anonymized form, and the next steps are planned through social learning tools.

[0346] As a concrete example, when user C is learning a language, the emotion engine detects from the user's voice that they are "satisfied." The server uses this emotion information to provide user C with the next level of advanced conversation scenarios, helping them improve their skills while maintaining their motivation.

[0347] Thus, the present invention allows users to receive an optimal learning experience tailored to their mood and state at any given time, dramatically improving the efficiency and satisfaction of language acquisition.

[0348] The following describes the processing flow.

[0349] Step 1:

[0350] The user starts a session on their device and enters their learning goal (e.g., conversational English for travel) and current language level. This information is necessary to create a customized learning plan to meet the user's individual needs.

[0351] Step 2:

[0352] The terminal sends user input information to the server. The server uses a learning plan generation mechanism to generate a specific learning plan based on the user's goals and level. This plan includes learning content and conversation scenarios.

[0353] Step 3:

[0354] The user provides voice input in response to a conversation scenario displayed on their device. The voice input receiving device captures this audio and sends it to the server.

[0355] Step 4:

[0356] The server processes the received audio using analysis tools and an emotion engine. During this process, the user's pronunciation and grammar are evaluated, while the emotion engine recognizes the user's emotional state (e.g., tension, joy, concentration) from the audio.

[0357] Step 5:

[0358] Based on the analysis results and recognized emotions, the server generates feedback. This feedback includes not only comments on pronunciation and grammar, but also encouragement and advice tailored to the user's emotions. For example, if tension is detected, simple exercises to relax may be recommended.

[0359] Step 6:

[0360] Through a presentation mechanism, the terminal presents feedback from the server to the user. The feedback is displayed as text or audio, which the user reviews and understands.

[0361] Step 7:

[0362] Based on the feedback, users can choose to continue with the next conversation practice or connect with other learners to utilize social learning methods. In this process, the user's emotional state is shared anonymously and used to facilitate interaction with other learners.

[0363] Step 8:

[0364] The server monitors the user's learning progress, adaptively adjusts the learning plan, and suggests new content and scenarios. This helps maintain the user's continuous learning and motivation.

[0365] This process enables the system to provide an optimal learning experience tailored to each user's individual needs, maximizing the effectiveness of language acquisition.

[0366] (Example 2)

[0367] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0368] In today's learning environment, there is a challenge in providing appropriate learning experiences tailored to the individual state and emotions of each learner. In particular, traditional learning methods do not adequately provide individualized feedback and adjustments based on the user's emotions and progress, and there is a need for efficient methods to maximize learning effectiveness.

[0369] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0370] In this invention, the server includes a plan generation means for generating a learning plan based on user input, an input receiving means for receiving voice input, an analysis means for analyzing the voice input and generating feedback, and an emotion analysis means for analyzing the user's emotional state and adjusting the learning experience. This makes it possible to provide the most appropriate learning experience for each individual user in real time.

[0371] The "plan generation means" is a function that generates an individualized learning plan based on the user's input information.

[0372] The "input receiving means" refers to a function that receives voice input from the user.

[0373] The "analysis means" refers to a function that analyzes received audio data and generates feedback related to pronunciation and grammar.

[0374] A "presentation method" is a function that provides the generated feedback to the user visually or audibly.

[0375] The "adjustment mechanism" is a function that dynamically adjusts the existing learning plan based on the user's learning progress.

[0376] "Emotional analysis tools" are functions that analyze the tone of the user's voice and word choices to recognize the user's emotional state.

[0377] This invention is a language learning system designed to enhance the user's individual learning experience. This system can generate dynamic learning plans tailored to the user's individual needs and adjust the learning experience based on emotions.

[0378] The user first accesses a language learning application through their device and inputs their learning goals and current proficiency level. The device sends this information to a server. The server uses a generative AI model to generate a personalized learning plan based on the provided data. This plan is then executed interactively via voice input.

[0379] The voice input from the user via the device is sent to the server by an input receiving device. The server analyzes the user's emotional state using an emotion analysis device and provides feedback and learning adjustments accordingly. A generative AI model provides optimal learning content based on the user's current learning status and emotions. Specifically, this includes providing a challenging task when the user is judged to be relaxed, or providing encouraging messages when the user is stressed.

[0380] The presentation methods allow users to visually or audibly confirm feedback and the next learning steps. This enables learners to progress at their own pace. Furthermore, by incorporating social learning methods that facilitate interaction with other learners, it is possible to further enrich the diversity of learning.

[0381] For example, if a user inputs "Teach me a new word" via voice, the server can select a word appropriate to its difficulty level and present it along with an example sentence. Another example of a prompt is, "Please use the emotion engine to provide appropriate feedback so that the user can learn in a relaxed state." In this way, it is possible to provide an optimal learning experience tailored to the user's emotions and state.

[0382] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0383] Step 1:

[0384] Users access the learning platform via their device and log in. After logging in, users enter their learning goals (e.g., mastering everyday conversation) and current proficiency level (e.g., beginner, intermediate, advanced). The entered data is sent from the device to the server and used as input to create a personalized learning plan.

[0385] Step 2:

[0386] The server uses a generative AI model to generate a personalized learning plan based on the user's learning goals and ability level. The generative AI model considers the user's past learning history and preference patterns to select the most suitable learning materials and assignments. This generates a learning plan that matches the user's needs and sends it to the device.

[0387] Step 3:

[0388] The user uses a device to input voice data and initiate an interactive learning session. The device receives the user's voice data in real time and sends it to the server. The voice data mainly consists of phrases and questions used during the learning process.

[0389] Step 4:

[0390] The server analyzes the audio data received from the terminal and generates feedback on pronunciation and grammar. The audio analysis also includes sentiment analysis, determining the user's emotions from their tone of voice and word choice. The generated feedback is adjusted according to the user's emotional state.

[0391] Step 5:

[0392] Based on the results of the emotion analysis, the server dynamically adjusts the user's learning experience. For example, if the server detects tension, it provides calming feedback and easier tasks; if it determines that the user is enjoying themselves, it offers more challenging tasks. The adjusted feedback is then sent from the server to the user's device.

[0393] Step 6:

[0394] The device presents the user with feedback and adjusted learning materials sent from the server. This is often done through text display or audio output. The user then uses this feedback to advance their learning.

[0395] Step 7:

[0396] If users wish to engage in further interaction or learning, they can repeat the same process. Furthermore, social learning features allow users to deepen their interactions with other learners, providing a more diverse learning experience.

[0397] (Application Example 2)

[0398] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0399] One challenge for users is that it's difficult to receive product recommendations that match their emotions and current mood when selecting products. This leads to information overload in the product selection process, causing users to struggle to make decisions.

[0400] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0401] In this invention, the server includes an emotion analysis device that analyzes the user's emotions and adjusts suggestions, a collaborative work device, and an information integration device. This makes it possible to provide optimal product information based on the user's emotions.

[0402] A "plan generation device" is a device that automatically creates an optimal learning and recommendation plan based on the user's input information.

[0403] An "input receiving device" is a device for receiving voice and text data provided by the user.

[0404] An "analysis device" is a device that analyzes received data and generates necessary information and feedback.

[0405] A "presentation device" is a device used to present feedback and information generated by an analysis device to the user.

[0406] A "adjustment device" is a device that optimizes and adjusts the plan according to the user's progress and emotions.

[0407] An "emotion analysis device" is a device that analyzes the user's current emotions from their voice or text and adjusts the operation of the entire system based on the results.

[0408] A "collaborative work device" is a device that provides an environment for performing tasks and learning in cooperation with other users.

[0409] An "information integration device" is a device that integrates various types of information related to the user and provides them in an appropriate format.

[0410] To implement this invention, an advanced interactive system is constructed to provide information and learning content tailored to the user's emotions. The system consists of a plan generation device, an input receiving device, an analysis device, a presentation device, an adjustment device, an emotion analysis device, a collaborative work device, and an information integration device.

[0411] The plan generation device generates learning plans and product recommendation plans that match the user's requests based on data entered by the user using a terminal. For example, suppose a user enters "I want to try a new language learning method" through a smartphone application. This input is sent to the server, received by the plan generation device, and provides a plan that suits the user's needs.

[0412] The input receiving device receives user voice and text input. In this system, for example, voice input is converted to text using the Google Speech-to-Text API. This converted text is then moved to the next parsing stage.

[0413] The analysis device analyzes the received data to determine pronunciation, grammar, and product review recommendation information.

[0414] Alternatively, it generates feedback. IBM Watson NLU, for example, is used for analysis.

[0415] The display device provides the user with visual or audible feedback on the results. This feedback is tailored to the user's preferences, such as presenting entertainment information for a visit to a restaurant with a strong Latin American theme.

[0416] The adjustment device fine-tunes the plan according to the user's progress and emotions. For example, if the user is feeling anxious, it will provide materials with a lower difficulty level.

[0417] The emotion analysis device reads the user's emotions from their voice and facial expressions. This is extremely useful in determining the system's subsequent behavior. The information presented is adjusted accordingly depending on whether the voice tone is tense or relaxed.

[0418] Collaborative work systems can link data when multiple users are viewing it simultaneously. For example, they can share opinions when a group of users are selecting a product together.

[0419] Information integration devices have the function of integrating information relevant to the user, thereby providing a highly satisfying experience. They retrieve the most relevant information from the database and construct a valuable experience for the user.

[0420] This design allows users to intuitively obtain the information they desire. An example of a prompt would be, "Show me reviews that will put my mind at ease." This prompt provides information that is sensitive to the user's emotional state.

[0421] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0422] Step 1:

[0423] The device receives voice input from the user. The user begins speaking into the device, and this is captured as audio data by the device's microphone. This audio data becomes the entry point for the system.

[0424] Step 2:

[0425] The device converts the captured audio data into text data using the Google Speech-to-Text API. The input here is audio data, and the output is text data. This converted text then proceeds to the next analysis phase.

[0426] Step 3:

[0427] The server analyzes text data using IBM Watson NLU to extract user emotions and intentions. The input is the text data generated in step 2, and the output is data on the user's emotions and specific intentions. This enables suggestions tailored to the user's emotions.

[0428] Step 4:

[0429] Based on the analyzed emotional information, the server's planning generator produces optimal recommendations and learning content for the user. The input here is emotional and intent data, and the output is a specific learning plan or product recommendation plan.

[0430] Step 5:

[0431] The server provides feedback to the user through a presentation device, which presents the generated plan. The information presented is communicated to the user in visual or audio format. The input is the plan generated in step 4, and the output is the feedback displayed or played back to the user.

[0432] Step 6:

[0433] Based on user feedback and usage history, the adjustment system fine-tunes the plan. The server uses this information to make suggestions more user-friendly for subsequent interactions. This cycle continuously improves the user experience.

[0434] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0435] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0436] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0437] [Third Embodiment]

[0438] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0439] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0440] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0441] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0442] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0443] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0444] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0445] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0446] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0447] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0448] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0449] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0450] This invention is a system that supports users' language learning and consists of a learning plan generation means, a voice input receiving means, an analysis means, a presentation means, an adjustment means, a social learning means, and a media integration means. Each means works in conjunction with the others to enable the provision of instruction tailored to the individual learning needs of the user.

[0451] First, the user inputs their language learning goals and current language level using a device. This information is sent from the device to the server. Based on this, the server generates a personalized learning plan using a learning plan generation mechanism. This learning plan includes necessary content and assignments and is dynamically adjusted according to the user's learning progress.

[0452] Next, the user responds verbally to the conversation scenario provided through the terminal. A voice input receiving device receives this voice input and sends it to the server. The server uses analysis tools, including voice analysis technology, to analyze the user's pronunciation and grammar in real time. The results of this analysis are generated as feedback regarding pronunciation accuracy and grammatical correctness.

[0453] The device presents this feedback to the user using a presentation method. The feedback is provided in text or audio format to aid the user's understanding.

[0454] Furthermore, social learning tools allow users to connect with other learners online and engage in collaborative language practice and discussions. This feature promotes interaction with learners worldwide and enhances social skills.

[0455] Furthermore, through media integration, the server selects and provides cultural background information, news articles, and videos related to the language the user is learning. This feature helps users learn the language in a practical context and contributes to deepening their cultural understanding.

[0456] As a concrete example, let's say User B is learning English with the goal of traveling. The terminal presents a conversation scenario at a check-in counter, and the user inputs, "Can I have my boarding pass, please?" The server analyzes the speech and provides appropriate response examples along with feedback on accurate pronunciation. In addition, by visually presenting videos about British culture and airport etiquette, the user can simultaneously acquire knowledge beyond pronunciation and grammar.

[0457] In this way, this system comprehensively supports language learning and effectively assists users in achieving their goals.

[0458] The following describes the processing flow.

[0459] Step 1:

[0460] Users launch a language learning application on their device and set individual language learning goals, such as "travel" or "business." They also select their current language level and take a self-assessment test to input their detailed skill level.

[0461] Step 2:

[0462] The terminal sends user input information to the server. This information includes learning objectives and language level assessment results.

[0463] Step 3:

[0464] Based on the received information, the server generates a user-specific learning plan using a learning plan generation mechanism. The generated plan, which includes specific conversation scenarios and related tasks, is sent to the user's terminal.

[0465] Step 4:

[0466] The user follows the instructions for pronunciation practice and makes voice input according to the conversation scenario displayed on the device. When the user reads a sample sentence aloud, the device sends the audio to the server.

[0467] Step 5:

[0468] The server receives voice input and uses analysis tools to evaluate pronunciation, grammar, and fluency. Speech recognition technology and AI-based natural language processing are used for the analysis, and evaluation results are generated.

[0469] Step 6:

[0470] The server generates and constructs feedback, which is then sent to the terminal. This feedback includes corrections to pronunciation and suggestions for improvement, helping the user learn.

[0471] Step 7:

[0472] The device displays the feedback it receives to the user. The feedback can be presented as text or audio, allowing the user to track their progress and move on to the next step.

[0473] Step 8:

[0474] Users can connect online with other learners through social learning tools. This allows users to practice conversation, share cultural information, and advance their learning further.

[0475] Step 9:

[0476] The server collects cultural information and the latest news related to the language the user is learning and transmits it to the terminal through a media integration system. This allows users to learn not only the language but also the culture behind it.

[0477] Step 10:

[0478] The user's learning progress is tracked over a certain period, and the server dynamically adjusts the learning plan based on this data, suggesting new challenges and learning materials. This improves the effectiveness and continuity of learning.

[0479] (Example 1)

[0480] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0481] Modern language education demands personalized learning plans and real-time analysis of speech input to provide feedback. However, existing systems are not sufficiently individualized, particularly in improving practical language application skills and cultural understanding. Furthermore, collaborative learning with other learners and the integration of relevant media have not been adequately achieved to enhance learning effectiveness. Therefore, the challenge lies in enabling users to effectively acquire language skills and apply them in real-world situations.

[0482] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0483] In this invention, the server includes a plan generation means, an audio data receiving means, and a data analysis means. This enables the generation of a personalized education plan that is dynamically adjusted according to the user's progress, real-time analysis of voice input, and provision of specific pronunciation and grammar correction instructions. Furthermore, data is sent and received effectively through secure communication means. In addition, collaborative learning means allow users to work with other educators and enhance their language skills in a practical setting. Moreover, media integration means allow users to gain a deeper understanding through relevant cultural background and media information.

[0484] "Plan generation means" refers to a function for creating personalized educational plans based on user input information.

[0485] "Audio data receiving means" refers to a function that receives audio information transmitted by the user on a terminal and sends it to a server.

[0486] "Data analysis means" refers to a function that instantly analyzes received audio information and evaluates the user's pronunciation and grammar.

[0487] "Information presentation means" refers to functions that present feedback generated by the server to the user to aid in understanding.

[0488] "Plan adjustment mechanism" refers to a function that dynamically modifies the educational plan according to the user's learning progress.

[0489] "Communication method" refers to the function of using protocols to securely send and receive data between a server and a terminal.

[0490] "Feedback generation means" refers to a function that evaluates the accuracy of pronunciation and grammar to the user based on the analyzed results and provides instructions.

[0491] "Collaborative learning tools" refer to features that allow users to collaborate with other educators and practice language together.

[0492] "Media integration means" refers to functions that integrate and provide cultural information and media content related to the language a user is learning.

[0493] This invention is a system for supporting users' language education, and it functions through the coordinated action of multiple components. Specific embodiments are described below.

[0494] First, the user enters their target language and current language level using a terminal. A dedicated application software runs on the terminal, accepting input through a user interface. The entered data is securely transmitted from the terminal to the server. This transmission utilizes encryption technologies such as the SSL / TLS protocol.

[0495] The server generates personalized learning plans using a plan generation mechanism based on the received user data. This generation process utilizes algorithms, for example, implemented in Python. The generated plans are dynamically adjusted based on the user's language learning progress. The learning plans include scenarios specifically tailored to different skills such as listening, speaking, and reading.

[0496] Next, the user uses their device to provide voice responses to a specified scenario. This voice is converted into a digital signal by the device and transmitted to the server via an audio data receiving device. The server instantly analyzes the voice data using an analysis device and evaluates the pronunciation and grammar. Specific voice analysis utilizes speech recognition APIs and natural language processing (NLP) technologies.

[0497] Based on the analysis results, the server uses a feedback generation mechanism to create a response for the user. This response includes accuracy of pronunciation and grammatical corrections. The terminal then uses an information presentation mechanism to present this feedback to the user in an easy-to-understand manner. The feedback is provided in both text and audio formats.

[0498] Furthermore, through collaborative learning methods, users can connect with other learners in real time and practice language together. Additionally, media integration allows the server to extract cultural background and media information and provide it to users through their terminals. This enables users to deepen not only their language acquisition but also their understanding of culture.

[0499] For example, if a user sets "Useful English Phrases for Travel" as their learning goal, the device will present a conversation scenario at an airport check-in counter. The user will say, "Can I have my boarding pass, please?" The server will analyze this audio and provide appropriate feedback. Furthermore, as relevant cultural background, a video about airport etiquette will be played on the device.

[0500] Examples of prompt messages are as follows:

[0501] "To help users practice everyday conversation, please set up situations and provide feedback on pronunciation and grammar."

[0502] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0503] Step 1:

[0504] The user launches the application on the terminal and enters the target language and current language level. During this process, text data is entered through the user interface, and the terminal converts it into a digital format. The entered information is securely transmitted to the server. The output is the user information data sent to the server.

[0505] Step 2:

[0506] The server analyzes the received user information and generates a personalized educational plan using a plan generation mechanism. An algorithm implemented in Python specifies appropriate tasks and content based on the user's goals. The input is digital data of the user information, and the output is the generated educational plan. This plan is stored on the server and sent to the terminal.

[0507] Step 3:

[0508] The user responds verbally to conversation scenarios provided via the terminal. The terminal uses a speech recognition device to convert the speech into a digital signal and sends this audio data to the server. The input is the user's voice data, and the output is the digital audio signal transferred to the server.

[0509] Step 4:

[0510] The server processes the received audio data using analysis tools. Specifically, it converts the audio to text using a speech recognition API and evaluates the grammar and pronunciation using natural language processing techniques. It generates feedback based on the analysis results. The input is a digital audio signal, and the output is the analysis results and the feedback data based on them.

[0511] Step 5:

[0512] The terminal presents the user with feedback received from the server. This feedback includes accuracy of pronunciation, grammatical corrections, and areas for improvement, and is presented to the user visually and audibly. The input is the feedback data, and the output is the feedback information presented to the user.

[0513] Step 6:

[0514] Users connect with other learners using collaborative learning tools to practice and discuss together. The device enables real-time interaction via an online platform. Input is communication data from other learners, and output is user interaction.

[0515] Step 7:

[0516] The server uses media integration to select cultural content relevant to the user's learning and provides it to the user through the terminal. This allows the user to learn language and related culture simultaneously. The input is data of the cultural content, and the output is media information presented to the user.

[0517] (Application Example 1)

[0518] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0519] Conventional language learning systems have struggled to provide flexible learning methods tailored to individual learning needs. Furthermore, they lacked sufficient mechanisms for learners to check and correct their pronunciation and grammatical errors in real time. Additionally, methods for learners to deepen their understanding of cultural context relevant to actual usage environments were limited. Therefore, the present invention aims to provide an individualized learning experience and improve users' multifaceted language abilities.

[0520] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0521] In this invention, the server includes planning means for generating a language learning plan based on user input, speech acquisition means for analyzing speech related to vocabulary training in real time, and material provision means for providing materials to the user through a visual medium. This allows the user to check and correct their pronunciation and grammar on the spot through always personalized, real-time feedback. It also facilitates understanding of cultural background through a visual medium.

[0522] A "planning mechanism" is a system for generating personalized language learning plans based on user input information.

[0523] A "voice acquisition means" is a mechanism for receiving and analyzing the voice emitted by the user in real time.

[0524] The "analysis means" is a mechanism for detecting pronunciation and grammatical errors based on received audio data and generating feedback.

[0525] A "display means" is a mechanism for presenting feedback generated by analysis to the user in text or audio format.

[0526] The "adjustment mechanism" is a system for updating and optimizing the generated learning plan in a timely manner according to the user's learning progress.

[0527] A "means of providing information" refers to a mechanism that provides users with materials including visual information and cultural background to promote a deeper understanding.

[0528] A "media integration tool" is a mechanism for integrating cultural information related to language learning and providing it to learners.

[0529] A "collaborative learning tool" is a mechanism that allows users to work together with other learners to train and practice language collaboratively.

[0530] The system for realizing this invention integrates a variety of functions to effectively support the user's language learning. Next, the operation of this system will be described.

[0531] Users first input their learning goals and current language level using a device such as a smartphone or tablet. The device then sends this information to the server. The server, as a planning tool, generates a personalized learning plan based on this user information. This plan can be flexibly modified according to the user's progress.

[0532] When a user practices speaking, the system receives the spoken audio in real time through the audio acquisition mechanism. The server processes this audio data using an analysis mechanism and generates feedback on pronunciation and grammatical errors. The display mechanism presents the feedback to the user to aid understanding.

[0533] Furthermore, the materials provided offer users visual information, such as cultural background and related video materials. This allows users to learn not only about language but also about the cultural aspects of the language in depth.

[0534] This system is implemented as a smartphone application using React Native, with AWS Lambda running on the cloud server. AWS Transcribe and PyDub are used for speech analysis. Feedback is synthesized using AWS Polly.

[0535] As a concrete example, a user might watch a scene from a movie and practice pronouncing the lines from it. In this case, the server analyzes the pronunciation and provides appropriate feedback in real time. Furthermore, it can provide the cultural background and related information of the scene as video material.

[0536] An example of a prompt sentence would be, "I want to watch a scene from a movie and try to pronounce it myself to get accurate feedback and cultural understanding."

[0537] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0538] Step 1:

[0539] The user inputs their learning goals and current language level from their device. The device sends this data to the server. The server receives the input data and uses a planning tool to generate a personalized learning plan tailored to the user. This plan includes the content and resources necessary for the user to achieve their immediate learning goals.

[0540] Step 2:

[0541] The user uses a device to perform voice input as part of language learning. The device captures voice data through its microphone and sends it to the server via a voice acquisition device. The server converts the received voice into text data using AWS Transcribe and performs detailed voice analysis using PyDub. Here, the waveform and phonemes of the voice data are analyzed to generate feedback on pronunciation and grammar.

[0542] Step 3:

[0543] The server sends the feedback obtained from the analysis to the terminal via a display device. The terminal displays this feedback on the screen in text format, or, if necessary, communicates it to the user in audio format using AWS Polly. The user can use this information to gain a better understanding of where improvements are needed.

[0544] Step 4:

[0545] The server uses a resource delivery system to select visual literature and videos relevant to the user's language learning and transmits this information to the terminal. The terminal displays this information to the user, facilitating an understanding of the language's cultural background and related topics. This allows the user to acquire a broad range of knowledge beyond mere language skills.

[0546] Step 5:

[0547] As users progress through their learning plan, they update their learning status based on new challenges and what they have learned. The server uses this feedback to periodically revise the learning plan using adjustment mechanisms, providing the user with the most optimal route. This enables flexible learning tailored to their progress.

[0548] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0549] This invention provides a system for more effective language learning, generating individualized learning plans for users and offering an interactive experience. In addition to learning plan generation means, voice input receiving means, analysis means, presentation means, and adjustment means, this system incorporates an emotion engine. This allows for the provision of a customized learning experience that takes user emotions into consideration.

[0550] First, the user accesses the application via their device and inputs their learning goals and current skill level. This generates a personalized learning plan. The device sends this information to the server, which then develops the learning plan.

[0551] The user initiates an interactive conversation by using voice input through their device. The emotion engine then analyzes the user's tone and word choice to recognize their emotions. This emotional information is used to further personalize the learning experience.

[0552] For example, if the server detects that the user is feeling stressed, it can provide relaxing scenarios or encouraging feedback to alleviate this. Conversely, if the user is enjoying themselves, the server can increase the difficulty level and provide more challenging tasks.

[0553] The generated learning feedback is presented to the user in audio or text format through various presentation methods. Furthermore, to facilitate interaction with other learners, the user's emotional state is used in an anonymized form, and the next steps are planned through social learning tools.

[0554] As a concrete example, when user C is learning a language, the emotion engine detects from the user's voice that they are "satisfied." The server uses this emotion information to provide user C with the next level of advanced conversation scenarios, helping them improve their skills while maintaining their motivation.

[0555] Thus, the present invention allows users to receive an optimal learning experience tailored to their mood and state at any given time, dramatically improving the efficiency and satisfaction of language acquisition.

[0556] The following describes the processing flow.

[0557] Step 1:

[0558] The user starts a session on their device and enters their learning goal (e.g., conversational English for travel) and current language level. This information is necessary to create a customized learning plan to meet the user's individual needs.

[0559] Step 2:

[0560] The terminal sends user input information to the server. The server uses a learning plan generation mechanism to generate a specific learning plan based on the user's goals and level. This plan includes learning content and conversation scenarios.

[0561] Step 3:

[0562] The user provides voice input in response to a conversation scenario displayed on their device. The voice input receiving device captures this audio and sends it to the server.

[0563] Step 4:

[0564] The server processes the received audio using analysis tools and an emotion engine. During this process, the user's pronunciation and grammar are evaluated, while the emotion engine recognizes the user's emotional state (e.g., tension, joy, concentration) from the audio.

[0565] Step 5:

[0566] Based on the analysis results and recognized emotions, the server generates feedback. This feedback includes not only comments on pronunciation and grammar, but also encouragement and advice tailored to the user's emotions. For example, if tension is detected, simple exercises to relax may be recommended.

[0567] Step 6:

[0568] Through a presentation mechanism, the terminal presents feedback from the server to the user. The feedback is displayed as text or audio, which the user reviews and understands.

[0569] Step 7:

[0570] Based on the feedback, users can choose to continue with the next conversation practice or connect with other learners to utilize social learning methods. In this process, the user's emotional state is shared anonymously and used to facilitate interaction with other learners.

[0571] Step 8:

[0572] The server monitors the user's learning progress, adaptively adjusts the learning plan, and suggests new content and scenarios. This helps maintain the user's continuous learning and motivation.

[0573] This process enables the system to provide an optimal learning experience tailored to each user's individual needs, maximizing the effectiveness of language acquisition.

[0574] (Example 2)

[0575] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0576] In today's learning environment, there is a challenge in providing appropriate learning experiences tailored to the individual state and emotions of each learner. In particular, traditional learning methods do not adequately provide individualized feedback and adjustments based on the user's emotions and progress, and there is a need for efficient methods to maximize learning effectiveness.

[0577] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0578] In this invention, the server includes a plan generation means for generating a learning plan based on user input, an input receiving means for receiving voice input, an analysis means for analyzing the voice input and generating feedback, and an emotion analysis means for analyzing the user's emotional state and adjusting the learning experience. This makes it possible to provide the most appropriate learning experience for each individual user in real time.

[0579] The "plan generation means" is a function that generates an individualized learning plan based on the user's input information.

[0580] The "input receiving means" refers to a function that receives voice input from the user.

[0581] The "analysis means" refers to a function that analyzes received audio data and generates feedback related to pronunciation and grammar.

[0582] A "presentation method" is a function that provides the generated feedback to the user visually or audibly.

[0583] The "adjustment mechanism" is a function that dynamically adjusts the existing learning plan based on the user's learning progress.

[0584] "Emotional analysis tools" are functions that analyze the tone of the user's voice and word choices to recognize the user's emotional state.

[0585] This invention is a language learning system designed to enhance the user's individual learning experience. This system can generate dynamic learning plans tailored to the user's individual needs and adjust the learning experience based on emotions.

[0586] The user first accesses a language learning application through their device and inputs their learning goals and current proficiency level. The device sends this information to a server. The server uses a generative AI model to generate a personalized learning plan based on the provided data. This plan is then executed interactively via voice input.

[0587] The voice input from the user via the device is sent to the server by an input receiving device. The server analyzes the user's emotional state using an emotion analysis device and provides feedback and learning adjustments accordingly. A generative AI model provides optimal learning content based on the user's current learning status and emotions. Specifically, this includes providing a challenging task when the user is judged to be relaxed, or providing encouraging messages when the user is stressed.

[0588] The presentation methods allow users to visually or audibly confirm feedback and the next learning steps. This enables learners to progress at their own pace. Furthermore, by incorporating social learning methods that facilitate interaction with other learners, it is possible to further enrich the diversity of learning.

[0589] For example, if a user inputs "Teach me a new word" via voice, the server can select a word appropriate to its difficulty level and present it along with an example sentence. Another example of a prompt is, "Please use the emotion engine to provide appropriate feedback so that the user can learn in a relaxed state." In this way, it is possible to provide an optimal learning experience tailored to the user's emotions and state.

[0590] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0591] Step 1:

[0592] Users access the learning platform via their device and log in. After logging in, users enter their learning goals (e.g., mastering everyday conversation) and current proficiency level (e.g., beginner, intermediate, advanced). The entered data is sent from the device to the server and used as input to create a personalized learning plan.

[0593] Step 2:

[0594] The server uses a generative AI model to generate a personalized learning plan based on the user's learning goals and ability level. The generative AI model considers the user's past learning history and preference patterns to select the most suitable learning materials and assignments. This generates a learning plan that matches the user's needs and sends it to the device.

[0595] Step 3:

[0596] The user uses a device to input voice data and initiate an interactive learning session. The device receives the user's voice data in real time and sends it to the server. The voice data mainly consists of phrases and questions used during the learning process.

[0597] Step 4:

[0598] The server analyzes the audio data received from the terminal and generates feedback on pronunciation and grammar. The audio analysis also includes sentiment analysis, determining the user's emotions from their tone of voice and word choice. The generated feedback is adjusted according to the user's emotional state.

[0599] Step 5:

[0600] Based on the results of the emotion analysis, the server dynamically adjusts the user's learning experience. For example, if the server detects tension, it provides calming feedback and easier tasks; if it determines that the user is enjoying themselves, it offers more challenging tasks. The adjusted feedback is then sent from the server to the user's device.

[0601] Step 6:

[0602] The device presents the user with feedback and adjusted learning materials sent from the server. This is often done through text display or audio output. The user then uses this feedback to advance their learning.

[0603] Step 7:

[0604] If users wish to engage in further interaction or learning, they can repeat the same process. Furthermore, social learning features allow users to deepen their interactions with other learners, providing a more diverse learning experience.

[0605] (Application Example 2)

[0606] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0607] One challenge for users is that it's difficult to receive product recommendations that match their emotions and current mood when selecting products. This leads to information overload in the product selection process, causing users to struggle to make decisions.

[0608] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0609] In this invention, the server includes an emotion analysis device that analyzes the user's emotions and adjusts suggestions, a collaborative work device, and an information integration device. This makes it possible to provide optimal product information based on the user's emotions.

[0610] A "plan generation device" is a device that automatically creates an optimal learning and recommendation plan based on the user's input information.

[0611] An "input receiving device" is a device for receiving voice and text data provided by the user.

[0612] An "analysis device" is a device that analyzes received data and generates necessary information and feedback.

[0613] A "presentation device" is a device used to present feedback and information generated by an analysis device to the user.

[0614] A "adjustment device" is a device that optimizes and adjusts the plan according to the user's progress and emotions.

[0615] An "emotion analysis device" is a device that analyzes the user's current emotions from their voice or text and adjusts the operation of the entire system based on the results.

[0616] A "collaborative work device" is a device that provides an environment for performing tasks and learning in cooperation with other users.

[0617] An "information integration device" is a device that integrates various types of information related to the user and provides them in an appropriate format.

[0618] To implement this invention, an advanced interactive system is constructed to provide information and learning content tailored to the user's emotions. The system consists of a plan generation device, an input receiving device, an analysis device, a presentation device, an adjustment device, an emotion analysis device, a collaborative work device, and an information integration device.

[0619] The plan generation device generates learning plans and product recommendation plans that match the user's requests based on data entered by the user using a terminal. For example, suppose a user enters "I want to try a new language learning method" through a smartphone application. This input is sent to the server, received by the plan generation device, and provides a plan that suits the user's needs.

[0620] The input receiving device receives user voice and text input. In this system, for example, voice input is converted to text using the Google Speech-to-Text API. This converted text is then moved to the next parsing stage.

[0621] The analysis device analyzes the received data to determine pronunciation, grammar, and product review recommendation information.

[0622] Alternatively, it generates feedback. IBM Watson NLU, for example, is used for analysis.

[0623] The display device provides the user with visual or audible feedback on the results. This feedback is tailored to the user's preferences, such as presenting entertainment information for a visit to a restaurant with a strong Latin American theme.

[0624] The adjustment device fine-tunes the plan according to the user's progress and emotions. For example, if the user is feeling anxious, it will provide materials with a lower difficulty level.

[0625] The emotion analysis device reads the user's emotions from their voice and facial expressions. This is extremely useful in determining the system's subsequent behavior. The information presented is adjusted accordingly depending on whether the voice tone is tense or relaxed.

[0626] Collaborative work systems can link data when multiple users are viewing it simultaneously. For example, they can share opinions when a group of users are selecting a product together.

[0627] Information integration devices have the function of integrating information relevant to the user, thereby providing a highly satisfying experience. They retrieve the most relevant information from the database and construct a valuable experience for the user.

[0628] This design allows users to intuitively obtain the information they desire. An example of a prompt would be, "Show me reviews that will put my mind at ease." This prompt provides information that is sensitive to the user's emotional state.

[0629] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0630] Step 1:

[0631] The device receives voice input from the user. The user begins speaking into the device, and this is captured as audio data by the device's microphone. This audio data becomes the entry point for the system.

[0632] Step 2:

[0633] The device converts the captured audio data into text data using the Google Speech-to-Text API. The input here is audio data, and the output is text data. This converted text then proceeds to the next analysis phase.

[0634] Step 3:

[0635] The server analyzes text data using IBM Watson NLU to extract user emotions and intentions. The input is the text data generated in step 2, and the output is data on the user's emotions and specific intentions. This enables suggestions tailored to the user's emotions.

[0636] Step 4:

[0637] Based on the analyzed emotional information, the server's planning generator produces optimal recommendations and learning content for the user. The input here is emotional and intent data, and the output is a specific learning plan or product recommendation plan.

[0638] Step 5:

[0639] The server provides feedback to the user through a presentation device, which presents the generated plan. The information presented is communicated to the user in visual or audio format. The input is the plan generated in step 4, and the output is the feedback displayed or played back to the user.

[0640] Step 6:

[0641] Based on user feedback and usage history, the adjustment system fine-tunes the plan. The server uses this information to make suggestions more user-friendly for subsequent interactions. This cycle continuously improves the user experience.

[0642] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0643] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0644] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0645] [Fourth Embodiment]

[0646] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0647] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0648] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0649] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0650] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0651] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0652] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0653] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0654] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0655] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0656] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0657] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0658] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0659] This invention is a system that supports users' language learning and consists of a learning plan generation means, a voice input receiving means, an analysis means, a presentation means, an adjustment means, a social learning means, and a media integration means. Each means works in conjunction with the others to enable the provision of instruction tailored to the individual learning needs of the user.

[0660] First, the user inputs their language learning goals and current language level using a device. This information is sent from the device to the server. Based on this, the server generates a personalized learning plan using a learning plan generation mechanism. This learning plan includes necessary content and assignments and is dynamically adjusted according to the user's learning progress.

[0661] Next, the user responds verbally to the conversation scenario provided through the terminal. A voice input receiving device receives this voice input and sends it to the server. The server uses analysis tools, including voice analysis technology, to analyze the user's pronunciation and grammar in real time. The results of this analysis are generated as feedback regarding pronunciation accuracy and grammatical correctness.

[0662] The device presents this feedback to the user using a presentation method. The feedback is provided in text or audio format to aid the user's understanding.

[0663] Furthermore, social learning tools allow users to connect with other learners online and engage in collaborative language practice and discussions. This feature promotes interaction with learners worldwide and enhances social skills.

[0664] Furthermore, through media integration, the server selects and provides cultural background information, news articles, and videos related to the language the user is learning. This feature helps users learn the language in a practical context and contributes to deepening their cultural understanding.

[0665] As a concrete example, let's say User B is learning English with the goal of traveling. The terminal presents a conversation scenario at a check-in counter, and the user inputs, "Can I have my boarding pass, please?" The server analyzes the speech and provides appropriate response examples along with feedback on accurate pronunciation. In addition, by visually presenting videos about British culture and airport etiquette, the user can simultaneously acquire knowledge beyond pronunciation and grammar.

[0666] In this way, this system comprehensively supports language learning and effectively assists users in achieving their goals.

[0667] The following describes the processing flow.

[0668] Step 1:

[0669] Users launch a language learning application on their device and set individual language learning goals, such as "travel" or "business." They also select their current language level and take a self-assessment test to input their detailed skill level.

[0670] Step 2:

[0671] The terminal sends user input information to the server. This information includes learning objectives and language level assessment results.

[0672] Step 3:

[0673] Based on the received information, the server generates a user-specific learning plan using a learning plan generation mechanism. The generated plan, which includes specific conversation scenarios and related tasks, is sent to the user's terminal.

[0674] Step 4:

[0675] The user follows the instructions for pronunciation practice and makes voice input according to the conversation scenario displayed on the device. When the user reads a sample sentence aloud, the device sends the audio to the server.

[0676] Step 5:

[0677] The server receives voice input and uses analysis tools to evaluate pronunciation, grammar, and fluency. Speech recognition technology and AI-based natural language processing are used for the analysis, and evaluation results are generated.

[0678] Step 6:

[0679] The server generates and constructs feedback, which is then sent to the terminal. This feedback includes corrections to pronunciation and suggestions for improvement, helping the user learn.

[0680] Step 7:

[0681] The device displays the feedback it receives to the user. The feedback can be presented as text or audio, allowing the user to track their progress and move on to the next step.

[0682] Step 8:

[0683] Users can connect online with other learners through social learning tools. This allows users to practice conversation, share cultural information, and advance their learning further.

[0684] Step 9:

[0685] The server collects cultural information and the latest news related to the language the user is learning and transmits it to the terminal through a media integration system. This allows users to learn not only the language but also the culture behind it.

[0686] Step 10:

[0687] The user's learning progress is tracked over a certain period, and the server dynamically adjusts the learning plan based on this data, suggesting new challenges and learning materials. This improves the effectiveness and continuity of learning.

[0688] (Example 1)

[0689] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0690] Modern language education demands personalized learning plans and real-time analysis of speech input to provide feedback. However, existing systems are not sufficiently individualized, particularly in improving practical language application skills and cultural understanding. Furthermore, collaborative learning with other learners and the integration of relevant media have not been adequately achieved to enhance learning effectiveness. Therefore, the challenge lies in enabling users to effectively acquire language skills and apply them in real-world situations.

[0691] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0692] In this invention, the server includes a plan generation means, an audio data receiving means, and a data analysis means. This enables the generation of a personalized education plan that is dynamically adjusted according to the user's progress, real-time analysis of voice input, and provision of specific pronunciation and grammar correction instructions. Furthermore, data is sent and received effectively through secure communication means. In addition, collaborative learning means allow users to work with other educators and enhance their language skills in a practical setting. Moreover, media integration means allow users to gain a deeper understanding through relevant cultural background and media information.

[0693] "Plan generation means" refers to a function for creating personalized educational plans based on user input information.

[0694] "Audio data receiving means" refers to a function that receives audio information transmitted by the user on a terminal and sends it to a server.

[0695] "Data analysis means" refers to a function that instantly analyzes received audio information and evaluates the user's pronunciation and grammar.

[0696] "Information presentation means" refers to functions that present feedback generated by the server to the user to aid in understanding.

[0697] "Plan adjustment mechanism" refers to a function that dynamically modifies the educational plan according to the user's learning progress.

[0698] "Communication method" refers to the function of using protocols to securely send and receive data between a server and a terminal.

[0699] "Feedback generation means" refers to a function that evaluates the accuracy of pronunciation and grammar to the user based on the analyzed results and provides instructions.

[0700] "Collaborative learning tools" refer to features that allow users to collaborate with other educators and practice language together.

[0701] "Media integration means" refers to functions that integrate and provide cultural information and media content related to the language a user is learning.

[0702] This invention is a system for supporting users' language education, and it functions through the coordinated action of multiple components. Specific embodiments are described below.

[0703] First, the user enters their target language and current language level using a terminal. A dedicated application software runs on the terminal, accepting input through a user interface. The entered data is securely transmitted from the terminal to the server. This transmission utilizes encryption technologies such as the SSL / TLS protocol.

[0704] The server generates personalized learning plans using a plan generation mechanism based on the received user data. This generation process utilizes algorithms, for example, implemented in Python. The generated plans are dynamically adjusted based on the user's language learning progress. The learning plans include scenarios specifically tailored to different skills such as listening, speaking, and reading.

[0705] Next, the user uses their device to provide voice responses to a specified scenario. This voice is converted into a digital signal by the device and transmitted to the server via an audio data receiving device. The server instantly analyzes the voice data using an analysis device and evaluates the pronunciation and grammar. Specific voice analysis utilizes speech recognition APIs and natural language processing (NLP) technologies.

[0706] Based on the analysis results, the server uses a feedback generation mechanism to create a response for the user. This response includes accuracy of pronunciation and grammatical corrections. The terminal then uses an information presentation mechanism to present this feedback to the user in an easy-to-understand manner. The feedback is provided in both text and audio formats.

[0707] Furthermore, through collaborative learning methods, users can connect with other learners in real time and practice language together. Additionally, media integration allows the server to extract cultural background and media information and provide it to users through their terminals. This enables users to deepen not only their language acquisition but also their understanding of culture.

[0708] For example, if a user sets "Useful English Phrases for Travel" as their learning goal, the device will present a conversation scenario at an airport check-in counter. The user will say, "Can I have my boarding pass, please?" The server will analyze this audio and provide appropriate feedback. Furthermore, as relevant cultural background, a video about airport etiquette will be played on the device.

[0709] Examples of prompt messages are as follows:

[0710] "To help users practice everyday conversation, please set up situations and provide feedback on pronunciation and grammar."

[0711] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0712] Step 1:

[0713] The user launches the application on the terminal and enters the target language and current language level. During this process, text data is entered through the user interface, and the terminal converts it into a digital format. The entered information is securely transmitted to the server. The output is the user information data sent to the server.

[0714] Step 2:

[0715] The server analyzes the received user information and generates a personalized educational plan using a plan generation mechanism. An algorithm implemented in Python specifies appropriate tasks and content based on the user's goals. The input is digital data of the user information, and the output is the generated educational plan. This plan is stored on the server and sent to the terminal.

[0716] Step 3:

[0717] The user responds verbally to conversation scenarios provided via the terminal. The terminal uses a speech recognition device to convert the speech into a digital signal and sends this audio data to the server. The input is the user's voice data, and the output is the digital audio signal transferred to the server.

[0718] Step 4:

[0719] The server processes the received audio data using analysis tools. Specifically, it converts the audio to text using a speech recognition API and evaluates the grammar and pronunciation using natural language processing techniques. It generates feedback based on the analysis results. The input is a digital audio signal, and the output is the analysis results and the feedback data based on them.

[0720] Step 5:

[0721] The terminal presents the user with feedback received from the server. This feedback includes accuracy of pronunciation, grammatical corrections, and areas for improvement, and is presented to the user visually and audibly. The input is the feedback data, and the output is the feedback information presented to the user.

[0722] Step 6:

[0723] Users connect with other learners using collaborative learning tools to practice and discuss together. The device enables real-time interaction via an online platform. Input is communication data from other learners, and output is user interaction.

[0724] Step 7:

[0725] The server uses media integration to select cultural content relevant to the user's learning and provides it to the user through the terminal. This allows the user to learn language and related culture simultaneously. The input is data of the cultural content, and the output is media information presented to the user.

[0726] (Application Example 1)

[0727] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0728] Conventional language learning systems have struggled to provide flexible learning methods tailored to individual learning needs. Furthermore, they lacked sufficient mechanisms for learners to check and correct their pronunciation and grammatical errors in real time. Additionally, methods for learners to deepen their understanding of cultural context relevant to actual usage environments were limited. Therefore, the present invention aims to provide an individualized learning experience and improve users' multifaceted language abilities.

[0729] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0730] In this invention, the server includes planning means for generating a language learning plan based on user input, speech acquisition means for analyzing speech related to vocabulary training in real time, and material provision means for providing materials to the user through a visual medium. This allows the user to check and correct their pronunciation and grammar on the spot through always personalized, real-time feedback. It also facilitates understanding of cultural background through a visual medium.

[0731] A "planning mechanism" is a system for generating personalized language learning plans based on user input information.

[0732] A "voice acquisition means" is a mechanism for receiving and analyzing the voice emitted by the user in real time.

[0733] The "analysis means" is a mechanism for detecting pronunciation and grammatical errors based on received audio data and generating feedback.

[0734] A "display means" is a mechanism for presenting feedback generated by analysis to the user in text or audio format.

[0735] The "adjustment mechanism" is a system for updating and optimizing the generated learning plan in a timely manner according to the user's learning progress.

[0736] A "means of providing information" refers to a mechanism that provides users with materials including visual information and cultural background to promote a deeper understanding.

[0737] A "media integration tool" is a mechanism for integrating cultural information related to language learning and providing it to learners.

[0738] A "collaborative learning tool" is a mechanism that allows users to work together with other learners to train and practice language collaboratively.

[0739] The system for realizing this invention integrates a variety of functions to effectively support the user's language learning. Next, the operation of this system will be described.

[0740] Users first input their learning goals and current language level using a device such as a smartphone or tablet. The device then sends this information to the server. The server, as a planning tool, generates a personalized learning plan based on this user information. This plan can be flexibly modified according to the user's progress.

[0741] When a user practices speaking, the system receives the spoken audio in real time through the audio acquisition mechanism. The server processes this audio data using an analysis mechanism and generates feedback on pronunciation and grammatical errors. The display mechanism presents the feedback to the user to aid understanding.

[0742] Furthermore, the materials provided offer users visual information, such as cultural background and related video materials. This allows users to learn not only about language but also about the cultural aspects of the language in depth.

[0743] This system is implemented as a smartphone application using React Native, with AWS Lambda running on the cloud server. AWS Transcribe and PyDub are used for speech analysis. Feedback is synthesized using AWS Polly.

[0744] As a concrete example, a user might watch a scene from a movie and practice pronouncing the lines from it. In this case, the server analyzes the pronunciation and provides appropriate feedback in real time. Furthermore, it can provide the cultural background and related information of the scene as video material.

[0745] An example of a prompt sentence would be, "I want to watch a scene from a movie and try to pronounce it myself to get accurate feedback and cultural understanding."

[0746] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0747] Step 1:

[0748] The user inputs their learning goals and current language level from their device. The device sends this data to the server. The server receives the input data and uses a planning tool to generate a personalized learning plan tailored to the user. This plan includes the content and resources necessary for the user to achieve their immediate learning goals.

[0749] Step 2:

[0750] The user uses a device to perform voice input as part of language learning. The device captures voice data through its microphone and sends it to the server via a voice acquisition device. The server converts the received voice into text data using AWS Transcribe and performs detailed voice analysis using PyDub. Here, the waveform and phonemes of the voice data are analyzed to generate feedback on pronunciation and grammar.

[0751] Step 3:

[0752] The server sends the feedback obtained from the analysis to the terminal via a display device. The terminal displays this feedback on the screen in text format, or, if necessary, communicates it to the user in audio format using AWS Polly. The user can use this information to gain a better understanding of where improvements are needed.

[0753] Step 4:

[0754] The server uses a resource delivery system to select visual literature and videos relevant to the user's language learning and transmits this information to the terminal. The terminal displays this information to the user, facilitating an understanding of the language's cultural background and related topics. This allows the user to acquire a broad range of knowledge beyond mere language skills.

[0755] Step 5:

[0756] As users progress through their learning plan, they update their learning status based on new challenges and what they have learned. The server uses this feedback to periodically revise the learning plan using adjustment mechanisms, providing the user with the most optimal route. This enables flexible learning tailored to their progress.

[0757] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0758] This invention provides a system for more effective language learning, generating individualized learning plans for users and offering an interactive experience. In addition to learning plan generation means, voice input receiving means, analysis means, presentation means, and adjustment means, this system incorporates an emotion engine. This allows for the provision of a customized learning experience that takes user emotions into consideration.

[0759] First, the user accesses the application via their device and inputs their learning goals and current skill level. This generates a personalized learning plan. The device sends this information to the server, which then develops the learning plan.

[0760] The user initiates an interactive conversation by using voice input through their device. The emotion engine then analyzes the user's tone and word choice to recognize their emotions. This emotional information is used to further personalize the learning experience.

[0761] For example, if the server detects that the user is feeling stressed, it can provide relaxing scenarios or encouraging feedback to alleviate this. Conversely, if the user is enjoying themselves, the server can increase the difficulty level and provide more challenging tasks.

[0762] The generated learning feedback is presented to the user in audio or text format through various presentation methods. Furthermore, to facilitate interaction with other learners, the user's emotional state is used in an anonymized form, and the next steps are planned through social learning tools.

[0763] As a concrete example, when user C is learning a language, the emotion engine detects from the user's voice that they are "satisfied." The server uses this emotion information to provide user C with the next level of advanced conversation scenarios, helping them improve their skills while maintaining their motivation.

[0764] Thus, the present invention allows users to receive an optimal learning experience tailored to their mood and state at any given time, dramatically improving the efficiency and satisfaction of language acquisition.

[0765] The following describes the processing flow.

[0766] Step 1:

[0767] The user starts a session on their device and enters their learning goal (e.g., conversational English for travel) and current language level. This information is necessary to create a customized learning plan to meet the user's individual needs.

[0768] Step 2:

[0769] The terminal sends user input information to the server. The server uses a learning plan generation mechanism to generate a specific learning plan based on the user's goals and level. This plan includes learning content and conversation scenarios.

[0770] Step 3:

[0771] The user provides voice input in response to a conversation scenario displayed on their device. The voice input receiving device captures this audio and sends it to the server.

[0772] Step 4:

[0773] The server processes the received audio using analysis tools and an emotion engine. During this process, the user's pronunciation and grammar are evaluated, while the emotion engine recognizes the user's emotional state (e.g., tension, joy, concentration) from the audio.

[0774] Step 5:

[0775] Based on the analysis results and recognized emotions, the server generates feedback. This feedback includes not only comments on pronunciation and grammar, but also encouragement and advice tailored to the user's emotions. For example, if tension is detected, simple exercises to relax may be recommended.

[0776] Step 6:

[0777] Through a presentation mechanism, the terminal presents feedback from the server to the user. The feedback is displayed as text or audio, which the user reviews and understands.

[0778] Step 7:

[0779] Based on the feedback, users can choose to continue with the next conversation practice or connect with other learners to utilize social learning methods. In this process, the user's emotional state is shared anonymously and used to facilitate interaction with other learners.

[0780] Step 8:

[0781] The server monitors the user's learning progress, adaptively adjusts the learning plan, and suggests new content and scenarios. This helps maintain the user's continuous learning and motivation.

[0782] This process enables the system to provide an optimal learning experience tailored to each user's individual needs, maximizing the effectiveness of language acquisition.

[0783] (Example 2)

[0784] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0785] In today's learning environment, there is a challenge in providing appropriate learning experiences tailored to the individual state and emotions of each learner. In particular, traditional learning methods do not adequately provide individualized feedback and adjustments based on the user's emotions and progress, and there is a need for efficient methods to maximize learning effectiveness.

[0786] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0787] In this invention, the server includes a plan generation means for generating a learning plan based on user input, an input receiving means for receiving voice input, an analysis means for analyzing the voice input and generating feedback, and an emotion analysis means for analyzing the user's emotional state and adjusting the learning experience. This makes it possible to provide the most appropriate learning experience for each individual user in real time.

[0788] The "plan generation means" is a function that generates an individualized learning plan based on the user's input information.

[0789] The "input receiving means" refers to a function that receives voice input from the user.

[0790] The "analysis means" refers to a function that analyzes received audio data and generates feedback related to pronunciation and grammar.

[0791] A "presentation method" is a function that provides the generated feedback to the user visually or audibly.

[0792] The "adjustment mechanism" is a function that dynamically adjusts the existing learning plan based on the user's learning progress.

[0793] "Emotional analysis tools" are functions that analyze the tone of the user's voice and word choices to recognize the user's emotional state.

[0794] This invention is a language learning system designed to enhance the user's individual learning experience. This system can generate dynamic learning plans tailored to the user's individual needs and adjust the learning experience based on emotions.

[0795] The user first accesses a language learning application through their device and inputs their learning goals and current proficiency level. The device sends this information to a server. The server uses a generative AI model to generate a personalized learning plan based on the provided data. This plan is then executed interactively via voice input.

[0796] The voice input from the user via the device is sent to the server by an input receiving device. The server analyzes the user's emotional state using an emotion analysis device and provides feedback and learning adjustments accordingly. A generative AI model provides optimal learning content based on the user's current learning status and emotions. Specifically, this includes providing a challenging task when the user is judged to be relaxed, or providing encouraging messages when the user is stressed.

[0797] The presentation methods allow users to visually or audibly confirm feedback and the next learning steps. This enables learners to progress at their own pace. Furthermore, by incorporating social learning methods that facilitate interaction with other learners, it is possible to further enrich the diversity of learning.

[0798] For example, if a user inputs "Teach me a new word" via voice, the server can select a word appropriate to its difficulty level and present it along with an example sentence. Another example of a prompt is, "Please use the emotion engine to provide appropriate feedback so that the user can learn in a relaxed state." In this way, it is possible to provide an optimal learning experience tailored to the user's emotions and state.

[0799] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0800] Step 1:

[0801] Users access the learning platform via their device and log in. After logging in, users enter their learning goals (e.g., mastering everyday conversation) and current proficiency level (e.g., beginner, intermediate, advanced). The entered data is sent from the device to the server and used as input to create a personalized learning plan.

[0802] Step 2:

[0803] The server uses a generative AI model to generate a personalized learning plan based on the user's learning goals and ability level. The generative AI model considers the user's past learning history and preference patterns to select the most suitable learning materials and assignments. This generates a learning plan that matches the user's needs and sends it to the device.

[0804] Step 3:

[0805] The user uses a device to input voice data and initiate an interactive learning session. The device receives the user's voice data in real time and sends it to the server. The voice data mainly consists of phrases and questions used during the learning process.

[0806] Step 4:

[0807] The server analyzes the audio data received from the terminal and generates feedback on pronunciation and grammar. The audio analysis also includes sentiment analysis, determining the user's emotions from their tone of voice and word choice. The generated feedback is adjusted according to the user's emotional state.

[0808] Step 5:

[0809] Based on the results of the emotion analysis, the server dynamically adjusts the user's learning experience. For example, if the server detects tension, it provides calming feedback and easier tasks; if it determines that the user is enjoying themselves, it offers more challenging tasks. The adjusted feedback is then sent from the server to the user's device.

[0810] Step 6:

[0811] The device presents the user with feedback and adjusted learning materials sent from the server. This is often done through text display or audio output. The user then uses this feedback to advance their learning.

[0812] Step 7:

[0813] If users wish to engage in further interaction or learning, they can repeat the same process. Furthermore, social learning features allow users to deepen their interactions with other learners, providing a more diverse learning experience.

[0814] (Application Example 2)

[0815] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0816] One challenge for users is that it's difficult to receive product recommendations that match their emotions and current mood when selecting products. This leads to information overload in the product selection process, causing users to struggle to make decisions.

[0817] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0818] In this invention, the server includes an emotion analysis device that analyzes the user's emotions and adjusts suggestions, a collaborative work device, and an information integration device. This makes it possible to provide optimal product information based on the user's emotions.

[0819] A "plan generation device" is a device that automatically creates an optimal learning and recommendation plan based on the user's input information.

[0820] An "input receiving device" is a device for receiving voice and text data provided by the user.

[0821] An "analysis device" is a device that analyzes received data and generates necessary information and feedback.

[0822] A "presentation device" is a device used to present feedback and information generated by an analysis device to the user.

[0823] A "adjustment device" is a device that optimizes and adjusts the plan according to the user's progress and emotions.

[0824] An "emotion analysis device" is a device that analyzes the user's current emotions from their voice or text and adjusts the operation of the entire system based on the results.

[0825] A "collaborative work device" is a device that provides an environment for performing tasks and learning in cooperation with other users.

[0826] An "information integration device" is a device that integrates various types of information related to the user and provides them in an appropriate format.

[0827] To implement this invention, an advanced interactive system is constructed to provide information and learning content tailored to the user's emotions. The system consists of a plan generation device, an input receiving device, an analysis device, a presentation device, an adjustment device, an emotion analysis device, a collaborative work device, and an information integration device.

[0828] The plan generation device generates learning plans and product recommendation plans that match the user's requests based on data entered by the user using a terminal. For example, suppose a user enters "I want to try a new language learning method" through a smartphone application. This input is sent to the server, received by the plan generation device, and provides a plan that suits the user's needs.

[0829] The input receiving device receives user voice and text input. In this system, for example, voice input is converted to text using the Google Speech-to-Text API. This converted text is then moved to the next parsing stage.

[0830] The analysis device analyzes the received data to determine pronunciation, grammar, and product review recommendation information.

[0831] Alternatively, it generates feedback. IBM Watson NLU, for example, is used for analysis.

[0832] The display device provides the user with visual or audible feedback on the results. This feedback is tailored to the user's preferences, such as presenting entertainment information for a visit to a restaurant with a strong Latin American theme.

[0833] The adjustment device fine-tunes the plan according to the user's progress and emotions. For example, if the user is feeling anxious, it will provide materials with a lower difficulty level.

[0834] The emotion analysis device reads the user's emotions from their voice and facial expressions. This is extremely useful in determining the system's subsequent behavior. The information presented is adjusted accordingly depending on whether the voice tone is tense or relaxed.

[0835] Collaborative work systems can link data when multiple users are viewing it simultaneously. For example, they can share opinions when a group of users are selecting a product together.

[0836] Information integration devices have the function of integrating information relevant to the user, thereby providing a highly satisfying experience. They retrieve the most relevant information from the database and construct a valuable experience for the user.

[0837] This design allows users to intuitively obtain the information they desire. An example of a prompt would be, "Show me reviews that will put my mind at ease." This prompt provides information that is sensitive to the user's emotional state.

[0838] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0839] Step 1:

[0840] The device receives voice input from the user. The user begins speaking into the device, and this is captured as audio data by the device's microphone. This audio data becomes the entry point for the system.

[0841] Step 2:

[0842] The device converts the captured audio data into text data using the Google Speech-to-Text API. The input here is audio data, and the output is text data. This converted text then proceeds to the next analysis phase.

[0843] Step 3:

[0844] The server analyzes text data using IBM Watson NLU to extract user emotions and intentions. The input is the text data generated in step 2, and the output is data on the user's emotions and specific intentions. This enables suggestions tailored to the user's emotions.

[0845] Step 4:

[0846] Based on the analyzed emotional information, the server's planning generator produces optimal recommendations and learning content for the user. The input here is emotional and intent data, and the output is a specific learning plan or product recommendation plan.

[0847] Step 5:

[0848] The server provides feedback to the user through a presentation device, which presents the generated plan. The information presented is communicated to the user in visual or audio format. The input is the plan generated in step 4, and the output is the feedback displayed or played back to the user.

[0849] Step 6:

[0850] Based on user feedback and usage history, the adjustment system fine-tunes the plan. The server uses this information to make suggestions more user-friendly for subsequent interactions. This cycle continuously improves the user experience.

[0851] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0852] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0853] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0854] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0855] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0856] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0857] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0858] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0859] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0860] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0861] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0862] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0863] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0864] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0865] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0866] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0867] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0868] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0869] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0870] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0871] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[0872] The following is further disclosed regarding the embodiments described above.

[0873] (Claim 1)

[0874] A learning plan generation means that generates a language learning plan based on user input,

[0875] A voice input receiving means that receives voice input corresponding to a conversation scenario provided in accordance with the learning plan,

[0876] An analysis means that analyzes the aforementioned voice input in real time and generates feedback regarding pronunciation and grammar,

[0877] A means for presenting the aforementioned feedback to the user,

[0878] An adjustment means for adjusting the learning plan based on the user's learning progress,

[0879] A system that includes this.

[0880] (Claim 2)

[0881] The system according to claim 1, comprising a social learning means for connecting users with other learners and conducting collaborative language practice.

[0882] (Claim 3)

[0883] The system according to claim 1, comprising media integration means for integrating cultural information and media content related to the language the user is learning.

[0884] "Example 1"

[0885] (Claim 1)

[0886] A plan generation means that generates an educational plan based on user input,

[0887] A sound data receiving means for receiving sound information corresponding to a scenario provided in accordance with the aforementioned educational plan,

[0888] A data analysis means that immediately analyzes the aforementioned sound data and generates a response regarding speech and grammar,

[0889] Information presentation means for presenting the aforementioned response to the user,

[0890] A plan adjustment means that adjusts the education plan based on the user's progress,

[0891] A means of communication for sending and receiving information using a secure communication protocol,

[0892] A feedback generation means that evaluates the user's pronunciation and grammar based on analyzed sound data and provides specific instructions for pronunciation accuracy and grammatical correction,

[0893] A system that includes this.

[0894] (Claim 2)

[0895] The system according to claim 1, comprising a means for collaborative learning to connect users with other educators and practice a language together.

[0896] (Claim 3)

[0897] The system according to claim 1, comprising media integration means for integrating cultural background and media information related to the language the user is learning.

[0898] "Application Example 1"

[0899] (Claim 1)

[0900] A planning means for generating a language learning plan based on user input,

[0901] A voice acquisition means that receives voice input related to vocabulary training provided in accordance with the learning plan,

[0902] An analysis means that analyzes the aforementioned voice input in real time and generates feedback regarding pronunciation and grammar,

[0903] A display means for presenting the aforementioned feedback to the user,

[0904] An adjustment means for adjusting the learning plan based on the user's learning progress,

[0905] A means of providing materials to users, including visual media, for performing audio analysis,

[0906] A system that includes this.

[0907] (Claim 2)

[0908] The system according to claim 1, comprising media integration means for integrating the user's cultural background information and visual materials to enhance the learning experience.

[0909] (Claim 3)

[0910] The system according to claim 1, comprising a collaborative learning means for connecting users with other learners and conducting language training together.

[0911] "Example 2 of combining an emotion engine"

[0912] (Claim 1)

[0913] A plan generation means that generates a learning plan based on user input,

[0914] An input receiving means for receiving voice input provided in accordance with the learning plan,

[0915] Analysis means for analyzing the aforementioned voice input and generating feedback regarding pronunciation and grammar,

[0916] A means for presenting the aforementioned feedback to the user,

[0917] An adjustment means for adjusting the learning plan based on the user's learning progress,

[0918] An emotion analysis method that analyzes the user's emotional state and adjusts the learning experience based on emotional information,

[0919] A system that includes this.

[0920] (Claim 2)

[0921] The system according to claim 1, comprising a social learning means for connecting users with other learners and conducting collaborative practice.

[0922] (Claim 3)

[0923] The system according to claim 1, comprising information integration means for integrating cultural information and media content related to the subject the user is learning.

[0924] "Application example 2 when combining with an emotional engine"

[0925] (Claim 1)

[0926] A plan generation device that generates a learning plan based on user input,

[0927] An input receiving device that receives voice input corresponding to a scenario provided in accordance with the learning plan,

[0928] An analysis device that analyzes the aforementioned voice input in real time and generates feedback regarding pronunciation and grammar,

[0929] A presentation device that presents the aforementioned feedback to the user,

[0930] An adjustment device that adjusts the learning plan based on the user's learning progress,

[0931] An emotion analysis device that analyzes the user's emotions and adjusts suggestions accordingly,

[0932] A system that includes this.

[0933] (Claim 2)

[0934] The system according to claim 1, comprising a collaborative work device for connecting users with other users and enabling them to work together.

[0935] (Claim 3)

[0936] The system according to claim 1, comprising an information integration device for integrating information related to services used by the user. [Explanation of Symbols]

[0937] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A learning plan generation means that generates a language learning plan based on user input, A voice input receiving means that receives voice input corresponding to a conversation scenario provided in accordance with the learning plan, An analysis means that analyzes the aforementioned voice input in real time and generates feedback regarding pronunciation and grammar, A means for presenting the aforementioned feedback to the user, An adjustment means for adjusting the learning plan based on the user's learning progress, A system that includes this.

2. The system according to claim 1, comprising a social learning means for connecting users with other learners and conducting collaborative language practice.

3. The system according to claim 1, comprising media integration means for integrating cultural information and media content related to the language the user is learning.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A