system

The system uses generative AI to create personalized language learning experiences through conversational interactions with a stuffed animal, addressing the lack of natural conversation in existing systems and enhancing learning effectiveness.

JP2026066722APending Publication Date: 2026-04-17SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-07
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing language learning systems fail to provide a natural conversation experience, limiting the learning effect.

Method used

A system comprising a lecture generation unit, output unit, collection unit, and question generation unit, utilizing generative AI to generate foreign language lectures and questions, outputting through a speaker attached to a stuffed animal, and collecting learner's voice for personalized feedback.

Benefits of technology

Enables effective language learning through natural conversational experiences, adapting content to learner's age, interests, and progress, providing real-time feedback on pronunciation and grammar.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026066722000001_ABST
    Figure 2026066722000001_ABST
Patent Text Reader

Abstract

The system according to this embodiment aims to effectively learn a language through a natural conversational experience. [Solution] The system according to the embodiment comprises a lecture generation unit, an output unit, a collection unit, a question generation unit, and a question output unit. The lecture generation unit generates audio of a lecture on a foreign language. The output unit outputs the audio generated by the lecture generation unit from a speaker attached to a stuffed animal. The collection unit collects the voices of the students. The question generation unit generates specific questions based on the audio collected by the collection unit. The question output unit outputs the audio of the questions generated by the question generation unit from a speaker.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the prior art, there is a problem that it is difficult to provide a natural conversation experience in language learning and the learning effect is limited.

[0005] The system according to the embodiment aims to effectively learn a language through a natural conversation experience.

Means for Solving the Problems

[0006] The system according to this embodiment comprises a lecture generation unit, an output unit, a collection unit, a question generation unit, and a question output unit. The lecture generation unit generates audio of a lecture on a foreign language. The output unit outputs the audio generated by the lecture generation unit through a speaker attached to a stuffed animal. The collection unit collects the voices of the students. The question generation unit generates specific questions based on the audio collected by the collection unit. The question output unit outputs the audio of the questions generated by the question generation unit through a speaker. [Effects of the Invention]

[0007] The system according to this embodiment allows for effective language learning through a natural conversational experience. [Brief explanation of the drawing]

[0008] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10]This shows an emotion map where multiple emotions are mapped. [Modes for carrying out the invention]

[0009] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0010] First, let's explain the terminology used in the following explanation.

[0011] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit).

[0012] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.

[0013] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0014] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor, an antenna, etc. The communication I / F manages communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.

[0016] [First Embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0017] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0019] The smart device 14 comprises a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The receiving device 38, output device 40, and camera 42 are also connected to the bus 52.

[0020] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, and accepts user input. The touch panel 38A accepts user input via touch by detecting contact with an object (e.g., a pen or finger). The microphone 38B accepts user input via voice by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 (see Figure 2) acquires the data indicating the user input.

[0021] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user by outputting the data in a form perceptible to the user (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0022] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0023] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0024] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0025] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0026] In the smart device 14, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart device 14 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0027] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device having the data generation model 58. The data processing device 12 may also be a server device or a terminal device owned by a user (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.

[0028] (Example of form 1) An embodiment of the present invention provides a language learning program that uses a stuffed animal with a generative AI directly or in a readily available configuration to enable children to learn natural pronunciation and conversational grammar. The system generates audio of a lecture on a foreign language and outputs this audio through a speaker attached to the stuffed animal. Next, it collects the learner's voice and generates appropriate questions based on that voice. The generated questions are also output from the speaker. This allows children to learn the language in a natural conversational setting. For example, if the stuffed animal says, "Hello, what did you do today?", the child might reply, "I went to school today." This voice is collected by the collection unit, and the AI ​​generates an appropriate question. For example, a question such as, "What did you study at school?" is generated and output from the speaker. This allows children to learn the language in a natural conversational setting. Furthermore, the question generation unit generates questions based on the learner's age, interests, and past conversational history. For example, if the child is interested in animals, a question such as, "What is your favorite animal?" is generated. It is also possible to evaluate the learner's pronunciation and grammar and generate questions based on the evaluation results. For example, if the child's pronunciation is inaccurate, a question such as, "Please try pronouncing it again," is generated. The question output unit includes a language selection unit to support multiple languages. This allows children to learn not only English but also other languages. For example, they can learn both English and Spanish. Furthermore, the question generation unit estimates the learner's emotions and generates questions based on those emotions. For example, if a child is tired, a question such as "Shall we take a short break?" is generated. It also includes a judgment unit that analyzes the audio of the answers to the questions output by the question output unit and determines the learner's level of learning based on the analysis results. This allows the lecture generation unit to adjust the difficulty level of the foreign language lectures based on the learner's level of learning. For example, if a child's level of learning is high, a more difficult lecture is generated. This allows the language learning program to enable children to learn languages ​​in a natural conversational setting. Note that the stuffed animals are robots that imitate humans or animals, but are not limited to such examples.

[0029] The language learning program according to this embodiment comprises a lecture generation unit, an output unit, a collection unit, a question generation unit, and a question output unit. The lecture generation unit generates audio of lectures on a foreign language. The lecture generation unit generates lectures on foreign language grammar and pronunciation, for example, using a generation AI. The generation AI can generate lecture content, for example, using a text generation AI (e.g., LLM). The lecture generation unit can also generate lectures on foreign language conversation practice using the generation AI. For example, the generation AI generates a conversation scenario and conducts a lecture based on it. The output unit outputs the audio generated by the lecture generation unit from a speaker attached to a stuffed animal. The output unit can output audio generated using the generation AI from the speaker, for example. The output unit can also adjust the tone and speed of the audio using the generation AI before outputting it. For example, the output unit adjusts the tone and speed of the audio according to the child's age and interests. The collection unit collects the voices of learners. The collection unit can collect the voices of learners, for example, using a microphone. Furthermore, the collection unit can analyze the collected audio using a generation AI. For example, the collection unit can analyze the accuracy of the learner's pronunciation and grammar. The question generation unit generates appropriate questions based on the audio collected by the collection unit. The question generation unit can, for example, use a generation AI to generate questions based on the accuracy of the learner's pronunciation and grammar. The question generation unit can also use a generation AI to generate questions based on the learner's age and interests. For example, if a child is interested in animals, the question generation unit will generate a question such as, "What is your favorite animal?" The question output unit outputs the audio of the questions generated by the question generation unit through the speaker. The question output unit can, for example, output the audio of questions generated using a generation AI through the speaker. The question output unit can also use a generation AI to adjust the tone and speed of the audio before outputting it. For example, the question output unit adjusts the tone and speed of the audio according to the child's age and interests. As a result, the language learning program according to this embodiment allows children to learn language in a natural conversational setting.

[0030] The lecture generation unit generates audio for lectures on foreign languages. Specifically, it uses a generative AI to generate lectures on foreign language grammar and pronunciation. The generative AI can generate lecture content using, for example, a text generation AI (e.g., LLM). The generative AI has learned from a vast amount of text data and can provide accurate information on grammar and pronunciation. For example, the generative AI can generate text explaining foreign language grammar rules and convert it into audio for lectures. The generative AI can also provide concrete examples of foreign language pronunciation to help students learn correct pronunciation. Furthermore, the lecture generation unit can also use the generative AI to generate lectures on foreign language conversation practice. For example, the generative AI can generate everyday conversation scenarios and conduct lectures based on them. The generative AI provides realistic conversation scenarios to help students learn phrases and expressions that can be used in actual conversations. This allows students to acquire skills that can be used in real conversations. Furthermore, the lecture generation unit can also use the generative AI to generate customized lectures tailored to the level and interests of the students. For example, beginner-level lectures focus on basic grammar and pronunciation, while advanced lectures provide more advanced grammar and pronunciation practice. It's also possible to generate lectures on specific topics based on the students' interests. This allows the lecture generation system to provide flexible lectures tailored to students' needs, supporting effective learning.

[0031] The output unit outputs the audio generated by the lecture generation unit through a speaker attached to the stuffed animal. Specifically, it can output audio generated using a generation AI through the speaker. The output unit can also adjust the tone and speed of the audio using the generation AI. For example, the output unit can adjust the tone and speed of the audio according to the child's age and interests. It can output the audio in a bright and friendly tone to make it more engaging for children. The output unit can also adjust the tone and speed of the audio in real time according to the learner's reactions. For example, it can output the audio at a slower speed to make it easier for learners to understand. Furthermore, the output unit can add appropriate emotional expressions depending on the content of the audio. For example, when asking a question, it can output the audio in an engaging tone, and when explaining something, it can output the audio in a calm tone. In this way, the output unit can make efforts to make the lecture content easier for learners to understand. Furthermore, the output unit can also use multiple speakers to produce three-dimensional audio output. This allows learners to have a more realistic lecture experience. For example, during conversation practice, different character voices can be output from different speakers, allowing learners to experience something closer to real conversation. This enables the output unit to provide a more effective learning environment for the learners.

[0032] The data collection unit collects the voices of the learners. Specifically, it can collect the learners' voices using a microphone. The data collection unit can also analyze the collected voices using generative AI. For example, the data collection unit can analyze the accuracy of the learners' pronunciation and grammar. The generative AI can analyze the learners' pronunciation using speech recognition technology and evaluate the accuracy of their pronunciation by comparing it to correct pronunciation. The generative AI can also analyze the learners' use of grammar and detect grammatical errors. This allows the data collection unit to accurately grasp the learners' learning progress and provide appropriate feedback. Furthermore, the data collection unit can also analyze the tone and speed of the learners' voices. For example, changes in tone and speed may occur when a learner is nervous or has insufficient understanding. The data collection unit can detect these changes and understand the learners' state. This allows the data collection unit to provide feedback that considers not only the learners' learning progress but also their psychological state. Furthermore, the data collection unit can accumulate learners' voice data and track their long-term learning progress. For example, the system can regularly evaluate the improvement in students' pronunciation and grammar, and confirm the effectiveness of their learning. This allows the data collection unit to continuously monitor students' learning progress and provide effective learning support.

[0033] The question generation unit generates appropriate questions based on the audio collected by the collection unit. Specifically, it can use a generation AI to generate questions based on the learner's pronunciation and grammatical accuracy. The generation AI can analyze the learner's learning progress and generate questions of appropriate difficulty. For example, if a learner has difficulty with a particular grammatical item, it can generate questions related to that grammatical item to help the learner deepen their understanding. The question generation unit can also use the generation AI to generate questions based on the learner's age and interests. For example, if a child is interested in animals, it can generate a question such as, "What is your favorite animal?" This allows the learner to learn while maintaining interest. Furthermore, the question generation unit can dynamically generate the next question in response to the learner's answer. For example, if a learner answers, "I like dogs," it can then generate a related question such as, "What do you like about dogs?" This allows the learner to learn within a natural conversational flow. In addition, the question generation unit can consider the learner's learning history and generate questions to review previously learned content. This helps the learner solidify what they have learned. The question generation unit can generate questions flexibly according to the learner's learning progress and interests, thereby supporting effective learning.

[0034] The question output unit outputs the audio of questions generated by the question generation unit through the speaker. Specifically, it can output the audio of questions generated using a generation AI through the speaker. The question output unit can also adjust the tone and speed of the audio using the generation AI. For example, the question output unit can adjust the tone and speed of the audio according to the child's age and interests. It can output the audio in a bright and friendly tone to make it easier for children to take an interest. Furthermore, the question output unit can adjust the tone and speed of the audio in real time according to the student's response. For example, it can output the audio at a slower speed to make it easier for students to understand. In addition, the question output unit can add appropriate emotional expressions depending on the content of the audio. For example, when asking a question, it can output the audio in an attention-grabbing tone, and when explaining, it can output the audio in a calm tone. In this way, the question output unit can make efforts to make the content of the questions easier for students to understand. Furthermore, the question output unit can also provide three-dimensional audio output using multiple speakers. This allows students to have a more realistic lecture experience. For example, during conversation practice, different characters' voices can be output from different speakers, allowing learners to experience something closer to a real conversation. This enables the question output unit to provide a more effective learning environment for the learners.

[0035] The question generation unit can generate questions based on the learner's age, interests, and past conversation history. For example, the question generation unit can generate questions of appropriate difficulty based on the learner's age. For example, it can generate simple questions for preschoolers and slightly more difficult questions for elementary school students. The question generation unit can also generate engaging questions based on the learner's interests. For example, it can generate a question such as, "What is your favorite animal?" for a child who is interested in animals. The question generation unit can also generate relevant questions based on past conversation history. For example, if the question generation unit was previously asked, "What did you study at school?", it can generate a question such as, "What classes did you have today?" based on that answer. This improves learning effectiveness by generating questions that are appropriate for the learner. Some or all of the above processing in the question generation unit may be performed using AI, for example, or not. For example, the question generation unit can generate questions using an AI model that takes the learner's age, interests, and past conversation history as input and outputs appropriate questions.

[0036] The question generation unit can evaluate the learner's pronunciation and grammar and generate questions based on the evaluation results. For example, the question generation unit can evaluate the accuracy of the learner's pronunciation and generate a question such as "Please try pronouncing it again" if the pronunciation is inaccurate. It can also evaluate the accuracy of the learner's grammar and generate a question such as "Please rephrase this sentence correctly" if the grammar is inaccurate. For example, the question generation unit generates questions that include appropriate feedback based on the evaluation results of the learner's pronunciation and grammar. This can promote improvement in the learner's pronunciation and grammar. Some or all of the above processing in the question generation unit may be performed using AI, for example, or without AI. For example, the question generation unit can generate questions using an AI model that takes the evaluation results of the learner's pronunciation and grammar as input and outputs appropriate questions.

[0037] The question output unit includes a language selection unit to support multiple languages. The question output unit can support multiple languages, such as English, French, and Spanish. The question output unit provides an interface for learners to select the language they wish to learn. For example, if a learner wants to learn English, the question output unit will output questions in English; if they want to learn Spanish, it will output questions in Spanish. The question output unit can also output questions in the appropriate language based on the learner's language selection. For example, if a learner wants to learn both English and Spanish, the question output unit will output questions in English and Spanish alternately. This allows learners to learn multiple languages. Some or all of the above processing in the question output unit may be performed using AI, for example, or without AI. For example, the question output unit can output questions using an AI model that takes the learner's language selection as input and outputs questions in the appropriate language.

[0038] The question output unit further includes a determination unit that analyzes the audio of the answers to the outputted questions and determines the learner's level of learning based on the analysis results, and the lecture generation unit can adjust the difficulty level of the foreign language lectures based on the learner's level of learning. For example, the question output unit collects and analyzes the learner's answer audio. Based on the analysis results, it evaluates the accuracy of the learner's pronunciation and grammar and determines the level of learning. For example, the question output unit determines a high level of learning if the learner's pronunciation is accurate and a low level of learning if the pronunciation is inaccurate. The question output unit can also determine a high level of learning if the learner's grammar is accurate and a low level of learning if the grammar is inaccurate. This improves learning effectiveness by providing lectures that are tailored to the learner's level of learning. Some or all of the above processing in the question output unit may be performed using AI, for example, or without using AI. For example, the question output unit can determine the level of learning using an AI model that takes the learner's answer audio as input and outputs the level of learning.

[0039] The lecture generation unit can analyze a student's past learning history and generate optimal lecture content. For example, the lecture generation unit can analyze a student's past learning history and generate lectures that focus on topics that need review or areas where the student struggles. For example, the lecture generation unit can generate lectures that review what the student has learned in the past. The lecture generation unit can also generate lectures that focus on areas where the student struggles. For example, if a student has difficulty with grammar, it will generate a lecture on grammar. The lecture generation unit can also generate lectures that include topics that the student is interested in. For example, if a student is interested in animals, it will generate a lecture that includes many examples related to animals. This improves learning effectiveness by providing lectures based on the student's past learning history. Some or all of the above processing in the lecture generation unit may be performed using AI, for example, or without AI. For example, the lecture generation unit can generate lectures using an AI model that takes the student's past learning history as input and outputs optimal lecture content.

[0040] The lecture generation unit can provide customized lectures based on the interests and concerns of the students. For example, the lecture generation unit generates lectures that include relevant topics based on the students' interests. For instance, if a student is interested in animals, the unit will generate lectures that include many examples related to animals. If a student is interested in sports, the unit can also generate lectures that include topics related to sports. For example, if a student is interested in music, the unit will generate lectures that include content related to music. This improves learning effectiveness by providing lectures tailored to the students' interests. Some or all of the above-described processes in the lecture generation unit may be performed using AI, for example, or without AI. For example, the lecture generation unit can generate lectures using an AI model that takes the students' interests as input and outputs customized lecture content.

[0041] The lecture generation unit can provide highly relevant content by considering the geographical location information of the students. For example, the lecture generation unit can generate lectures that include topics related to the region, taking into account the students' geographical location information. For example, if a student lives in a specific region, it can generate lectures that include culture and geographical features related to that region. If a student is traveling, it can also generate lectures that include content related to their travel destination. For example, if a student belongs to a specific cultural area, it can generate lectures that include content related to that culture. By providing lectures based on the students' geographical location information, the learning effect is improved. Some or all of the above processing in the lecture generation unit may be performed using AI, for example, or without AI. For example, the lecture generation unit can generate lectures using an AI model that takes the students' geographical location information as input and outputs highly relevant lecture content.

[0042] The lecture generation unit can analyze the social media activity of students and include relevant topics. For example, the lecture generation unit can analyze the social media activity of students and generate lectures that include frequently mentioned topics. For example, it can generate lectures that include content related to topics that students frequently mention on social media. It can also generate lectures related to the content of accounts that students follow. For example, it can generate lectures that include content related to topics in online communities that students participate in. This improves learning effectiveness by providing lectures based on students' social media activity. Some or all of the above processing in the lecture generation unit may be performed using AI, for example, or not using AI. For example, the lecture generation unit can generate lectures using an AI model that takes students' social media activity as input and outputs lecture content that includes relevant topics.

[0043] The output unit can adjust the frequency and volume of the audio based on the learner's auditory characteristics. For example, the output unit outputs audio at an appropriate frequency and volume, taking into account the learner's auditory characteristics. For instance, if the learner has difficulty hearing high frequencies, the output unit outputs audio at a lower frequency, and if the learner has difficulty hearing low frequencies, it outputs audio at a higher frequency. Furthermore, if the learner has a hearing impairment, the output unit can output audio at an appropriate volume. This improves learning effectiveness by providing audio output tailored to the learner's auditory characteristics. Some or all of the above processing in the output unit may be performed using AI, for example, or without AI. For example, the output unit can output audio using an AI model that takes the learner's auditory characteristics as input and outputs appropriate frequencies and volumes.

[0044] The output unit can detect the learner's ambient sounds and provide optimal audio output. For example, the output unit can detect ambient sounds around the learner and output audio at an appropriate volume and sound quality. For example, if the learner is in a noisy environment, the output unit will increase the volume of the audio output, and if the learner is in a quiet environment, it will decrease the volume of the audio output. It can also consider wind noise when the learner is outdoors when outputting audio. This improves learning effectiveness by providing audio output that is tailored to the learner's ambient sounds. Some or all of the above processing in the output unit may be performed using AI, for example, or without AI. For example, the output unit can output audio using an AI model that takes the learner's ambient sounds as input and outputs audio at an optimal volume and sound quality.

[0045] The output unit can select the optimal audio format considering the learner's device information. For example, the output unit can select an appropriate audio format based on the type of device the learner is using. For instance, if the learner is using a smartphone, the output unit will select the audio format best suited for smartphones; if they are using a tablet, it will select the audio format best suited for tablets. It can also select the audio format best suited for PCs if the learner is using a PC. This improves learning effectiveness by providing audio output tailored to the learner's device information. Some or all of the above processing in the output unit may be performed using AI, for example, or without AI. For example, the output unit can output audio using an AI model that takes the learner's device information as input and outputs the optimal audio format.

[0046] The output unit can analyze the learner's past responses and produce optimal audio output. For example, the output unit analyzes the learner's past responses and outputs audio based on preferred voice tone and speed. For example, the output unit can output optimal audio based on the voice tone the learner has preferred in the past. It can also output optimal audio based on the voice speed to which the learner has responded well in the past. For example, the output unit can output optimal audio based on the voice intonation the learner has preferred in the past. This improves learning effectiveness by providing audio output based on the learner's past responses. Some or all of the above processing in the output unit may be performed using AI, for example, or without AI. For example, the output unit can output audio using an AI model that takes the learner's past responses as input and outputs optimal audio.

[0047] The data collection unit can evaluate the accuracy of learners' pronunciation and grammar in real time. For example, the data collection unit can evaluate the accuracy of learners' pronunciation and grammar in real time and provide appropriate feedback. For example, if a learner's pronunciation is inaccurate, the data collection unit can provide real-time feedback, and if their grammar is inaccurate, it can suggest corrections in real time. It can also praise learners in real time if their pronunciation and grammar are accurate. This improves learning effectiveness by evaluating the accuracy of learners' pronunciation and grammar in real time. Some or all of the above processing in the data collection unit may be performed using AI, for example, or without AI. For example, the data collection unit can perform evaluation using an AI model that takes the accuracy of learners' pronunciation and grammar as input, evaluates it in real time, and outputs feedback.

[0048] The audio collection unit can filter out background noise from the learner to collect clear audio. For example, the audio collection unit filters out background noise from the learner's surroundings to remove noise and collect clear audio. For example, if the learner is in a noisy environment, the audio collection unit filters out background noise to collect audio, and if the learner is in a quiet environment, it minimizes background noise to collect audio. It can also filter out wind noise if the learner is outdoors. In this way, clear audio can be collected by filtering out background noise. Some or all of the above processing in the audio collection unit may be performed using AI, for example, or without AI. For example, the audio collection unit can collect audio using an AI model that takes the learner's background noise as input and outputs clear audio by removing noise.

[0049] The collection unit can prioritize collecting audio that is highly relevant, taking into account the learner's geographical location. For example, the collection unit can prioritize collecting audio related to a region, taking into account the learner's geographical location. For instance, if the learner lives in a specific region, the collection unit will prioritize collecting audio related to that region; if the learner is traveling, it will prioritize collecting audio related to their travel destination. Furthermore, if the learner belongs to a specific cultural sphere, it can prioritize collecting audio related to that culture. This improves learning effectiveness by prioritizing the collection of audio based on the learner's geographical location. Some or all of the processing described above in the collection unit may be performed using AI, for example, or without AI. For example, the collection unit can collect audio using an AI model that takes the learner's geographical location as input and outputs highly relevant audio.

[0050] The data collection unit can analyze the learner's social media activity and collect relevant audio. For example, the data collection unit can analyze the learner's social media activity and collect audio related to frequently mentioned topics. For example, the data collection unit can collect audio related to topics that the learner frequently mentions on social media and audio related to the content of accounts that the learner follows. It can also collect audio related to topics in online communities that the learner participates in. This improves learning effectiveness by collecting audio based on the learner's social media activity. Some or all of the processing described above in the data collection unit may be performed using AI, for example, or not. For example, the data collection unit can collect audio using an AI model that takes the learner's social media activity as input and outputs relevant audio.

[0051] The question generation unit can analyze the learner's past answer history and generate optimal questions. For example, the question generation unit can analyze the learner's past answer history and generate relevant questions. For example, the question generation unit can generate relevant questions based on questions the learner has answered in the past. It can also generate supplementary questions based on questions the learner has struggled with in the past. For example, the question generation unit can generate engaging questions based on topics the learner has shown interest in in the past. This improves learning effectiveness by providing questions based on the learner's past answer history. Some or all of the above processing in the question generation unit may be performed using AI, for example, or without AI. For example, the question generation unit can generate questions using an AI model that takes the learner's past answer history as input and outputs the optimal question.

[0052] The question generation unit can provide customized questions based on the learner's interests and concerns. For example, the question generation unit generates relevant questions based on the learner's interests and concerns. For instance, if the learner is interested in animals, the question generation unit will generate questions about animals; if the learner is interested in sports, it will generate questions related to sports. It can also generate questions about music if the learner is interested in music. This improves learning effectiveness by providing questions based on the learner's interests and concerns. Some or all of the above processing in the question generation unit may be performed using AI, for example, or without AI. For example, the question generation unit can generate questions using an AI model that takes the learner's interests and concerns as input and outputs customized questions.

[0053] The question generation unit can provide highly relevant questions by taking into account the learner's geographical location. For example, the question generation unit can generate questions related to the learner's region by considering their geographical location. For instance, if the learner lives in a specific region, the question generation unit can generate questions related to that region; if the learner is traveling, it can generate questions related to their travel destination. It can also generate questions related to a specific culture if the learner belongs to that culture. This improves learning effectiveness by providing questions based on the learner's geographical location. Some or all of the above processing in the question generation unit may be performed using AI, for example, or without AI. For example, the question generation unit can generate questions using an AI model that takes the learner's geographical location as input and outputs highly relevant questions.

[0054] The question generation unit can analyze the learner's social media activity and include relevant questions. For example, the question generation unit can analyze the learner's social media activity and generate questions related to frequently mentioned topics. For example, the question generation unit can generate questions related to topics that the learner frequently mentions on social media and questions related to the content of accounts that the learner follows. It can also generate questions related to topics in online communities that the learner participates in. This improves learning effectiveness by providing questions based on the learner's social media activity. Some or all of the above processing in the question generation unit may be performed using AI, for example, or not using AI. For example, the question generation unit can generate questions using an AI model that takes the learner's social media activity as input and outputs relevant questions.

[0055] The question output unit can adjust the frequency and volume of the audio based on the learner's auditory characteristics. For example, the question output unit outputs questions at an appropriate frequency and volume, taking into account the learner's auditory characteristics. For instance, if the learner has difficulty hearing high frequencies, the question output unit outputs questions at a low frequency, and if the learner has difficulty hearing low frequencies, it outputs questions at a high frequency. Furthermore, if the learner has a hearing impairment, the question output unit can output questions at an appropriate volume. This improves learning effectiveness by providing audio output of questions tailored to the learner's auditory characteristics. Some or all of the above processing in the question output unit may be performed using AI, for example, or without AI. For example, the question output unit can output questions using an AI model that takes the learner's auditory characteristics as input and outputs questions at an appropriate frequency and volume.

[0056] The question output unit can detect the learner's ambient sounds and output the question at an optimal volume and quality. For example, the question output unit can detect ambient sounds around the learner and output the question at an appropriate volume and quality. For example, if the learner is in a noisy environment, the question output unit will increase the volume of the question and output the question at a lower volume if the learner is in a quiet environment. It can also consider wind noise when the learner is outdoors when outputting the question. This improves learning effectiveness by providing audio output of questions that are tailored to the learner's ambient sounds. Some or all of the above processing in the question output unit may be performed using AI, for example, or without AI. For example, the question output unit can output questions using an AI model that takes the learner's ambient sounds as input and outputs questions at an optimal volume and quality.

[0057] The question output unit can select the optimal audio format considering the learner's device information. For example, the question output unit can select an appropriate audio format based on the type of device the learner is using. For instance, if the learner is using a smartphone, the question output unit will select the audio format best suited for smartphones; if they are using a tablet, it will select the audio format best suited for tablets. It can also select the audio format best suited for PCs if the learner is using a PC. This improves learning effectiveness by providing audio output of questions tailored to the learner's device information. Some or all of the above processing in the question output unit may be performed using AI, for example, or without AI. For example, the question output unit can output questions using an AI model that takes the learner's device information as input and outputs questions in the optimal audio format.

[0058] The question output unit can analyze the learner's past responses and output the optimal voice. For example, the question output unit analyzes the learner's past responses and outputs questions based on preferred voice tone and speed. For example, the question output unit outputs questions in the optimal voice based on the voice tone the learner has preferred in the past. It can also output questions in the optimal voice based on the voice speed the learner has responded well to in the past. For example, the question output unit outputs questions in the optimal voice based on the voice intonation the learner has preferred in the past. This improves learning effectiveness by providing voice output of questions based on the learner's past responses. Some or all of the above processing in the question output unit may be performed using AI, for example, or without AI. For example, the question output unit can output questions using an AI model that takes the learner's past responses as input and outputs questions in the optimal voice.

[0059] The language selection unit can analyze the learner's past learning history and suggest the most suitable language. For example, the language selection unit can analyze the learner's past learning history and suggest relevant languages. For example, the language selection unit can suggest relevant languages ​​based on the languages ​​the learner has learned in the past. It can also suggest supplementary languages ​​based on languages ​​the learner has struggled with in the past. For example, the language selection unit can suggest languages ​​that will interest the learner based on languages ​​the learner has shown interest in in the past. This improves learning effectiveness by providing language selection suggestions based on the learner's past learning history. Some or all of the above processing in the language selection unit may be performed using AI, for example, or without AI. For example, the language selection unit can make language selection suggestions using an AI model that takes the learner's past learning history as input and outputs the most suitable language.

[0060] The language selection unit can suggest highly relevant languages ​​by considering the learner's geographical location. For example, the language selection unit can suggest the language used in a particular region, taking into account the learner's geographical location. For example, if the learner lives in a specific region, the language selection unit can suggest the language used in that region; if the learner is traveling, it can suggest the language used in their travel destination. It can also suggest the language used in a particular culture if the learner belongs to that culture. This improves learning effectiveness by providing language selection suggestions based on the learner's geographical location. Some or all of the above processing in the language selection unit may be performed using AI, for example, or without AI. For example, the language selection unit can suggest language selections using an AI model that takes the learner's geographical location as input and outputs highly relevant languages.

[0061] The evaluation unit can analyze the learner's past learning history and set optimal evaluation criteria. For example, the evaluation unit can analyze the learner's past learning history and set appropriate evaluation criteria. For example, the evaluation unit can set appropriate evaluation criteria based on what the learner has learned in the past. It can also set supplementary evaluation criteria based on what the learner has struggled with in the past. For example, the evaluation unit can set attractive evaluation criteria based on what the learner has shown interest in in the past. This improves learning effectiveness by providing evaluation criteria based on the learner's past learning history. Some or all of the above processing in the evaluation unit may be performed using AI, for example, or without AI. For example, the evaluation unit can set evaluation criteria using an AI model that takes the learner's past learning history as input and outputs optimal evaluation criteria.

[0062] The evaluation unit can set highly relevant evaluation criteria by considering the learner's geographical location information. For example, the evaluation unit can set evaluation criteria used in a particular region by considering the learner's geographical location information. For example, if the learner lives in a particular region, the evaluation unit can set evaluation criteria used in that region, and if the learner is traveling, it can set evaluation criteria used in the travel destination. It can also set evaluation criteria used in a particular culture if the learner belongs to a particular cultural area. This improves learning effectiveness by providing evaluation criteria based on the learner's geographical location information. Some or all of the above processing in the evaluation unit may be performed using AI, for example, or without using AI. For example, the evaluation unit can set evaluation criteria using an AI model that takes the learner's geographical location information as input and outputs highly relevant evaluation criteria.

[0063] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0064] Language learning programs can provide customized lectures based on the learner's learning style. For example, they can generate lectures that heavily utilize images and videos for visual learners, and lectures that heavily utilize audio and music for auditory learners. They can also generate lectures that include interactive elements for tactile learners. By providing lectures tailored to the learner's learning style, learning effectiveness is improved. The lecture generation unit can generate lectures using an AI model that takes the learner's learning style as input and outputs the most suitable lecture content.

[0065] Language learning programs can provide real-time feedback based on the learner's progress. For example, if a learner completes a specific task, they can immediately receive positive feedback; if they are struggling with a task, they can receive additional hints and support. Furthermore, once a learner achieves a certain level of progress, guidance can be provided to help them move on to the next step. This improves learning effectiveness by providing feedback tailored to the learner's progress. The feedback system uses an AI model that takes the learner's progress as input and outputs appropriate feedback.

[0066] Language learning programs can provide customized learning plans based on the learner's learning objectives. For example, if a learner aims to pass a specific exam, a learning plan tailored to that exam can be provided; if their goal is to acquire conversational skills, a learning plan focusing on practical conversation practice can be provided. Furthermore, if a learner wants to learn business English, a learning plan tailored to business situations can be provided. This improves learning effectiveness by providing learning plans that match the learner's objectives. The learning plan generation unit can generate learning plans using an AI model that takes the learner's learning objectives as input and outputs the optimal learning plan.

[0067] The language learning program can suggest the optimal learning method based on the learner's learning environment. For example, if a learner is studying at home, it can suggest a learning method that takes place in a quiet environment; if a learner is studying during their commute, it can suggest a method that allows them to study while on the go. Furthermore, if a learner is studying in a group, it can suggest a learning method that incorporates group discussions and collaborative work. This improves learning effectiveness by providing learning methods tailored to the learner's environment. The learning method suggestion unit can suggest learning methods using an AI model that takes the learner's learning environment as input and outputs the optimal learning method.

[0068] The language learning program can suggest review timings based on the learner's learning history. For example, if a certain period has passed since the learner studied a particular topic, it can suggest a review of that topic to reinforce what they have learned. It can also suggest regular reviews of topics that the learner has struggled with in the past. By providing review timings based on the learner's learning history, the learning effect is improved. The review suggestion unit can suggest review timings using an AI model that takes the learner's learning history as input and outputs the optimal review timing.

[0069] The following briefly describes the processing flow for example form 1.

[0070] Step 1: The lecture generation unit generates audio for lectures on foreign languages. For example, it uses a generation AI to generate lectures on foreign language grammar and pronunciation. The generation AI can generate lecture content using a text generation AI (e.g., LLM). It can also use the generation AI to generate lectures on foreign language conversation practice. For example, the generation AI generates a conversation scenario and delivers a lecture based on it. Step 2: The output unit outputs the audio generated by the lecture generation unit through a speaker attached to the stuffed animal. For example, audio generated using a generation AI can be output through the speaker. The output unit can also use the generation AI to adjust the tone and speed of the audio before outputting it. For example, the tone and speed of the audio can be adjusted according to the child's age and interests. Step 3: The collection unit collects the participants' voices. For example, it can collect participants' voices using a microphone. It can also analyze the collected audio using generative AI. For example, it can analyze the accuracy of the participants' pronunciation and grammar. Step 4: The question generation unit generates appropriate questions based on the audio collected by the collection unit. For example, using a generation AI, questions can be generated based on the learner's pronunciation and grammatical accuracy. Alternatively, the generation AI can be used to generate questions based on the learner's age and interests. For example, if a child is interested in animals, the AI ​​can generate a question such as, "What is your favorite animal?" Step 5: The question output unit outputs the audio of the question generated by the question generation unit through the speaker. For example, the audio of a question generated using a generation AI can be output through the speaker. It is also possible to adjust the tone and speed of the audio using the generation AI. For example, the tone and speed of the audio can be adjusted according to the child's age and interests.

[0071] (Example of form 2) An embodiment of the present invention provides a language learning program that uses a stuffed animal with a generative AI directly or in a readily available configuration to enable children to learn natural pronunciation and conversational grammar. The system generates audio of a lecture on a foreign language and outputs this audio through a speaker attached to the stuffed animal. Next, it collects the learner's voice and generates appropriate questions based on that voice. The generated questions are also output from the speaker. This allows children to learn the language in a natural conversational setting. For example, if the stuffed animal says, "Hello, what did you do today?", the child might reply, "I went to school today." This voice is collected by the collection unit, and the AI ​​generates an appropriate question. For example, a question such as, "What did you study at school?" is generated and output from the speaker. This allows children to learn the language in a natural conversational setting. Furthermore, the question generation unit generates questions based on the learner's age, interests, and past conversational history. For example, if the child is interested in animals, a question such as, "What is your favorite animal?" is generated. It is also possible to evaluate the learner's pronunciation and grammar and generate questions based on the evaluation results. For example, if the child's pronunciation is inaccurate, a question such as, "Please try pronouncing it again," is generated. The question output unit includes a language selection unit to support multiple languages. This allows children to learn not only English but also other languages. For example, they can learn both English and Spanish. Furthermore, the question generation unit estimates the learner's emotions and generates questions based on those emotions. For example, if a child is tired, a question such as "Shall we take a short break?" is generated. It also includes a judgment unit that analyzes the audio of the answers to the questions output by the question output unit and determines the learner's level of learning based on the analysis results. This allows the lecture generation unit to adjust the difficulty level of the foreign language lectures based on the learner's level of learning. For example, if a child's level of learning is high, a more difficult lecture is generated. This allows the language learning program to enable children to learn languages ​​in a natural conversational setting. Note that the stuffed animals are robots that imitate humans or animals, but are not limited to such examples.

[0072] The language learning program according to this embodiment comprises a lecture generation unit, an output unit, a collection unit, a question generation unit, and a question output unit. The lecture generation unit generates audio of lectures on a foreign language. The lecture generation unit generates lectures on foreign language grammar and pronunciation, for example, using a generation AI. The generation AI can generate lecture content, for example, using a text generation AI (e.g., LLM). The lecture generation unit can also generate lectures on foreign language conversation practice using the generation AI. For example, the generation AI generates a conversation scenario and conducts a lecture based on it. The output unit outputs the audio generated by the lecture generation unit from a speaker attached to a stuffed animal. The output unit can output audio generated using the generation AI from the speaker, for example. The output unit can also adjust the tone and speed of the audio using the generation AI before outputting it. For example, the output unit adjusts the tone and speed of the audio according to the child's age and interests. The collection unit collects the voices of learners. The collection unit can collect the voices of learners, for example, using a microphone. Furthermore, the collection unit can analyze the collected audio using a generation AI. For example, the collection unit can analyze the accuracy of the learner's pronunciation and grammar. The question generation unit generates appropriate questions based on the audio collected by the collection unit. The question generation unit can, for example, use a generation AI to generate questions based on the accuracy of the learner's pronunciation and grammar. The question generation unit can also use a generation AI to generate questions based on the learner's age and interests. For example, if a child is interested in animals, the question generation unit will generate a question such as, "What is your favorite animal?" The question output unit outputs the audio of the questions generated by the question generation unit through the speaker. The question output unit can, for example, output the audio of questions generated using a generation AI through the speaker. The question output unit can also use a generation AI to adjust the tone and speed of the audio before outputting it. For example, the question output unit adjusts the tone and speed of the audio according to the child's age and interests. As a result, the language learning program according to this embodiment allows children to learn language in a natural conversational setting.

[0073] The lecture generation unit generates audio for lectures on foreign languages. Specifically, it uses a generative AI to generate lectures on foreign language grammar and pronunciation. The generative AI can generate lecture content using, for example, a text generation AI (e.g., LLM). The generative AI has learned from a vast amount of text data and can provide accurate information on grammar and pronunciation. For example, the generative AI can generate text explaining foreign language grammar rules and convert it into audio for lectures. The generative AI can also provide concrete examples of foreign language pronunciation to help students learn correct pronunciation. Furthermore, the lecture generation unit can also use the generative AI to generate lectures on foreign language conversation practice. For example, the generative AI can generate everyday conversation scenarios and conduct lectures based on them. The generative AI provides realistic conversation scenarios to help students learn phrases and expressions that can be used in actual conversations. This allows students to acquire skills that can be used in real conversations. Furthermore, the lecture generation unit can also use the generative AI to generate customized lectures tailored to the level and interests of the students. For example, beginner-level lectures focus on basic grammar and pronunciation, while advanced lectures provide more advanced grammar and pronunciation practice. It's also possible to generate lectures on specific topics based on the students' interests. This allows the lecture generation system to provide flexible lectures tailored to students' needs, supporting effective learning.

[0074] The output unit outputs the audio generated by the lecture generation unit through a speaker attached to the stuffed animal. Specifically, it can output audio generated using a generation AI through the speaker. The output unit can also adjust the tone and speed of the audio using the generation AI. For example, the output unit can adjust the tone and speed of the audio according to the child's age and interests. It can output the audio in a bright and friendly tone to make it more engaging for children. The output unit can also adjust the tone and speed of the audio in real time according to the learner's reactions. For example, it can output the audio at a slower speed to make it easier for learners to understand. Furthermore, the output unit can add appropriate emotional expressions depending on the content of the audio. For example, when asking a question, it can output the audio in an engaging tone, and when explaining something, it can output the audio in a calm tone. In this way, the output unit can make efforts to make the lecture content easier for learners to understand. Furthermore, the output unit can also use multiple speakers to produce three-dimensional audio output. This allows learners to have a more realistic lecture experience. For example, during conversation practice, different character voices can be output from different speakers, allowing learners to experience something closer to real conversation. This enables the output unit to provide a more effective learning environment for the learners.

[0075] The data collection unit collects the voices of the learners. Specifically, it can collect the learners' voices using a microphone. The data collection unit can also analyze the collected voices using generative AI. For example, the data collection unit can analyze the accuracy of the learners' pronunciation and grammar. The generative AI can analyze the learners' pronunciation using speech recognition technology and evaluate the accuracy of their pronunciation by comparing it to correct pronunciation. The generative AI can also analyze the learners' use of grammar and detect grammatical errors. This allows the data collection unit to accurately grasp the learners' learning progress and provide appropriate feedback. Furthermore, the data collection unit can also analyze the tone and speed of the learners' voices. For example, changes in tone and speed may occur when a learner is nervous or has insufficient understanding. The data collection unit can detect these changes and understand the learners' state. This allows the data collection unit to provide feedback that considers not only the learners' learning progress but also their psychological state. Furthermore, the data collection unit can accumulate learners' voice data and track their long-term learning progress. For example, the system can regularly evaluate the improvement in students' pronunciation and grammar, and confirm the effectiveness of their learning. This allows the data collection unit to continuously monitor students' learning progress and provide effective learning support.

[0076] The question generation unit generates appropriate questions based on the audio collected by the collection unit. Specifically, it can use a generation AI to generate questions based on the learner's pronunciation and grammatical accuracy. The generation AI can analyze the learner's learning progress and generate questions of appropriate difficulty. For example, if a learner has difficulty with a particular grammatical item, it can generate questions related to that grammatical item to help the learner deepen their understanding. The question generation unit can also use the generation AI to generate questions based on the learner's age and interests. For example, if a child is interested in animals, it can generate a question such as, "What is your favorite animal?" This allows the learner to learn while maintaining interest. Furthermore, the question generation unit can dynamically generate the next question in response to the learner's answer. For example, if a learner answers, "I like dogs," it can then generate a related question such as, "What do you like about dogs?" This allows the learner to learn within a natural conversational flow. In addition, the question generation unit can consider the learner's learning history and generate questions to review previously learned content. This helps the learner solidify what they have learned. The question generation unit can generate questions flexibly according to the learner's learning progress and interests, thereby supporting effective learning.

[0077] The question output unit outputs the audio of questions generated by the question generation unit through the speaker. Specifically, it can output the audio of questions generated using a generation AI through the speaker. The question output unit can also adjust the tone and speed of the audio using the generation AI. For example, the question output unit can adjust the tone and speed of the audio according to the child's age and interests. It can output the audio in a bright and friendly tone to make it easier for children to take an interest. Furthermore, the question output unit can adjust the tone and speed of the audio in real time according to the student's response. For example, it can output the audio at a slower speed to make it easier for students to understand. In addition, the question output unit can add appropriate emotional expressions depending on the content of the audio. For example, when asking a question, it can output the audio in an attention-grabbing tone, and when explaining, it can output the audio in a calm tone. In this way, the question output unit can make efforts to make the content of the questions easier for students to understand. Furthermore, the question output unit can also provide three-dimensional audio output using multiple speakers. This allows students to have a more realistic lecture experience. For example, during conversation practice, different characters' voices can be output from different speakers, allowing learners to experience something closer to a real conversation. This enables the question output unit to provide a more effective learning environment for the learners.

[0078] The question generation unit can generate questions based on the learner's age, interests, and past conversation history. For example, the question generation unit can generate questions of appropriate difficulty based on the learner's age. For example, it can generate simple questions for preschoolers and slightly more difficult questions for elementary school students. The question generation unit can also generate engaging questions based on the learner's interests. For example, it can generate a question such as, "What is your favorite animal?" for a child who is interested in animals. The question generation unit can also generate relevant questions based on past conversation history. For example, if the question generation unit was previously asked, "What did you study at school?", it can generate a question such as, "What classes did you have today?" based on that answer. This improves learning effectiveness by generating questions that are appropriate for the learner. Some or all of the above processing in the question generation unit may be performed using AI, for example, or not. For example, the question generation unit can generate questions using an AI model that takes the learner's age, interests, and past conversation history as input and outputs appropriate questions.

[0079] The question generation unit can evaluate the learner's pronunciation and grammar and generate questions based on the evaluation results. For example, the question generation unit can evaluate the accuracy of the learner's pronunciation and generate a question such as "Please try pronouncing it again" if the pronunciation is inaccurate. It can also evaluate the accuracy of the learner's grammar and generate a question such as "Please rephrase this sentence correctly" if the grammar is inaccurate. For example, the question generation unit generates questions that include appropriate feedback based on the evaluation results of the learner's pronunciation and grammar. This can promote improvement in the learner's pronunciation and grammar. Some or all of the above processing in the question generation unit may be performed using AI, for example, or without AI. For example, the question generation unit can generate questions using an AI model that takes the evaluation results of the learner's pronunciation and grammar as input and outputs appropriate questions.

[0080] The question output unit includes a language selection unit to support multiple languages. The question output unit can support multiple languages, such as English, French, and Spanish. The question output unit provides an interface for learners to select the language they wish to learn. For example, if a learner wants to learn English, the question output unit will output questions in English; if they want to learn Spanish, it will output questions in Spanish. The question output unit can also output questions in the appropriate language based on the learner's language selection. For example, if a learner wants to learn both English and Spanish, the question output unit will output questions in English and Spanish alternately. This allows learners to learn multiple languages. Some or all of the above processing in the question output unit may be performed using AI, for example, or without AI. For example, the question output unit can output questions using an AI model that takes the learner's language selection as input and outputs questions in the appropriate language.

[0081] The question generation unit can estimate the learner's emotions and generate questions based on those estimated emotions. For example, the question generation unit can estimate emotions from the learner's facial expressions and voice and generate questions appropriate to those emotions. For example, if the learner is tired, the question generation unit may generate a question such as, "Shall we take a short break?". It can also generate a question such as, "Did something fun happen?" if the learner is excited. For example, the question generation unit generates questions with appropriate tone and content based on the learner's emotions. This improves learning effectiveness by generating questions that are appropriate to the learner's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or a generative AI. The generative AI may be, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above-described processes in the question generation unit may be performed using AI, or not. For example, the question generation unit can generate questions using an AI model that takes the learner's facial expressions and voice data as input, estimates emotions, and outputs appropriate questions.

[0082] The question output unit further includes a determination unit that analyzes the audio of the answers to the outputted questions and determines the learner's level of learning based on the analysis results, and the lecture generation unit can adjust the difficulty level of the foreign language lectures based on the learner's level of learning. For example, the question output unit collects and analyzes the learner's answer audio. Based on the analysis results, it evaluates the accuracy of the learner's pronunciation and grammar and determines the level of learning. For example, the question output unit determines a high level of learning if the learner's pronunciation is accurate and a low level of learning if the pronunciation is inaccurate. The question output unit can also determine a high level of learning if the learner's grammar is accurate and a low level of learning if the grammar is inaccurate. This improves learning effectiveness by providing lectures that are tailored to the learner's level of learning. Some or all of the above processing in the question output unit may be performed using AI, for example, or without using AI. For example, the question output unit can determine the level of learning using an AI model that takes the learner's answer audio as input and outputs the level of learning.

[0083] The lecture generation unit can estimate the emotions of the learners and adjust the content and tone of the lecture based on the estimated emotions. For example, the lecture generation unit can estimate emotions from the learners' facial expressions and voice and adjust the lecture content and tone accordingly. For example, if the learners are excited, the lecture generation unit will proceed with the lecture in a calm tone, and if they are tired, it will include relaxing content. It can also add more difficult content if the learners are concentrating. This improves learning effectiveness by providing lectures that are tailored to the learners' emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or a generative AI. The generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the lecture generation unit may be performed using AI, for example, or without AI. For example, the lecture generation unit can generate a lecture using an AI model that takes learners' facial expressions and voice data as input, estimates emotions, and outputs appropriate lecture content and tone.

[0084] The lecture generation unit can analyze a student's past learning history and generate optimal lecture content. For example, the lecture generation unit can analyze a student's past learning history and generate lectures that focus on topics that need review or areas where the student struggles. For example, the lecture generation unit can generate lectures that review what the student has learned in the past. The lecture generation unit can also generate lectures that focus on areas where the student struggles. For example, if a student has difficulty with grammar, it will generate a lecture on grammar. The lecture generation unit can also generate lectures that include topics that the student is interested in. For example, if a student is interested in animals, it will generate a lecture that includes many examples related to animals. This improves learning effectiveness by providing lectures based on the student's past learning history. Some or all of the above processing in the lecture generation unit may be performed using AI, for example, or without AI. For example, the lecture generation unit can generate lectures using an AI model that takes the student's past learning history as input and outputs optimal lecture content.

[0085] The lecture generation unit can provide customized lectures based on the interests and concerns of the students. For example, the lecture generation unit generates lectures that include relevant topics based on the students' interests. For instance, if a student is interested in animals, the unit will generate lectures that include many examples related to animals. If a student is interested in sports, the unit can also generate lectures that include topics related to sports. For example, if a student is interested in music, the unit will generate lectures that include content related to music. This improves learning effectiveness by providing lectures tailored to the students' interests. Some or all of the above-described processes in the lecture generation unit may be performed using AI, for example, or without AI. For example, the lecture generation unit can generate lectures using an AI model that takes the students' interests as input and outputs customized lecture content.

[0086] The lecture generation unit can estimate the emotions of the learners and adjust the length of the lecture based on the estimated emotions. For example, the lecture generation unit can estimate emotions from the learners' facial expressions and voice and adjust the length of the lecture accordingly. For example, the lecture generation unit can provide a shorter lecture when the learners are tired and a longer lecture when they are focused. It can also provide a lecture of an appropriate length when the learners are excited. By providing lecture lengths that match the learners' emotions, the learning effect is improved. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or a generative AI. The generative AI is a text generation AI (e.g., LLM) or a multimodal generation AI, but is not limited to these examples. Some or all of the above processing in the lecture generation unit may be performed using AI, for example, or without AI. For example, the lecture generation unit can generate lectures using an AI model that takes learners' facial expressions and voice data as input, estimates emotions, and outputs an appropriate lecture length.

[0087] The lecture generation unit can provide highly relevant content by considering the geographical location information of the students. For example, the lecture generation unit can generate lectures that include topics related to the region, taking into account the students' geographical location information. For example, if a student lives in a specific region, it can generate lectures that include culture and geographical features related to that region. If a student is traveling, it can also generate lectures that include content related to their travel destination. For example, if a student belongs to a specific cultural area, it can generate lectures that include content related to that culture. By providing lectures based on the students' geographical location information, the learning effect is improved. Some or all of the above processing in the lecture generation unit may be performed using AI, for example, or without AI. For example, the lecture generation unit can generate lectures using an AI model that takes the students' geographical location information as input and outputs highly relevant lecture content.

[0088] The lecture generation unit can analyze the social media activity of students and include relevant topics. For example, the lecture generation unit can analyze the social media activity of students and generate lectures that include frequently mentioned topics. For example, it can generate lectures that include content related to topics that students frequently mention on social media. It can also generate lectures related to the content of accounts that students follow. For example, it can generate lectures that include content related to topics in online communities that students participate in. This improves learning effectiveness by providing lectures based on students' social media activity. Some or all of the above processing in the lecture generation unit may be performed using AI, for example, or not using AI. For example, the lecture generation unit can generate lectures using an AI model that takes students' social media activity as input and outputs lecture content that includes relevant topics.

[0089] The output unit can estimate the learner's emotions and adjust the tone and speed of the audio based on the estimated emotions. For example, the output unit can estimate emotions from the learner's facial expressions and voice and adjust the tone and speed of the audio accordingly. For example, the output unit can output audio in a relaxed tone when the learner is relaxed and at a fast speed when the learner is in a hurry. It can also output audio in a bright tone when the learner is excited. This improves learning effectiveness by providing audio output that matches the learner's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the output unit may be performed using AI, or not using AI. For example, the output unit can output audio using an AI model that takes the learner's facial expressions and voice data as input, estimates emotions, and outputs appropriate tone and speed of audio.

[0090] The output unit can adjust the frequency and volume of the audio based on the learner's auditory characteristics. For example, the output unit outputs audio at an appropriate frequency and volume, taking into account the learner's auditory characteristics. For instance, if the learner has difficulty hearing high frequencies, the output unit outputs audio at a lower frequency, and if the learner has difficulty hearing low frequencies, it outputs audio at a higher frequency. Furthermore, if the learner has a hearing impairment, the output unit can output audio at an appropriate volume. This improves learning effectiveness by providing audio output tailored to the learner's auditory characteristics. Some or all of the above processing in the output unit may be performed using AI, for example, or without AI. For example, the output unit can output audio using an AI model that takes the learner's auditory characteristics as input and outputs appropriate frequencies and volumes.

[0091] The output unit can detect the learner's ambient sounds and provide optimal audio output. For example, the output unit can detect ambient sounds around the learner and output audio at an appropriate volume and sound quality. For example, if the learner is in a noisy environment, the output unit will increase the volume of the audio output, and if the learner is in a quiet environment, it will decrease the volume of the audio output. It can also consider wind noise when the learner is outdoors when outputting audio. This improves learning effectiveness by providing audio output that is tailored to the learner's ambient sounds. Some or all of the above processing in the output unit may be performed using AI, for example, or without AI. For example, the output unit can output audio using an AI model that takes the learner's ambient sounds as input and outputs audio at an optimal volume and sound quality.

[0092] The output unit can estimate the learner's emotions and adjust the intonation of the voice based on the estimated emotions. For example, the output unit can estimate emotions from the learner's facial expressions and voice and adjust the intonation of the voice according to the emotion. For example, the output unit can output voice with a calm intonation when the learner is relaxed and with an energetic intonation when the learner is excited. It can also output voice with a calm intonation when the learner is tired. This improves learning effectiveness by providing voice output that matches the learner's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or a generative AI. The generative AI is, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above processing in the output unit may be performed using AI, for example, or without AI. For example, the output unit can output voice using an AI model that takes the learner's facial expressions and voice data as input, estimates emotions, and outputs appropriate intonation.

[0093] The output unit can select the optimal audio format considering the learner's device information. For example, the output unit can select an appropriate audio format based on the type of device the learner is using. For instance, if the learner is using a smartphone, the output unit will select the audio format best suited for smartphones; if they are using a tablet, it will select the audio format best suited for tablets. It can also select the audio format best suited for PCs if the learner is using a PC. This improves learning effectiveness by providing audio output tailored to the learner's device information. Some or all of the above processing in the output unit may be performed using AI, for example, or without AI. For example, the output unit can output audio using an AI model that takes the learner's device information as input and outputs the optimal audio format.

[0094] The output unit can analyze the learner's past responses and produce optimal audio output. For example, the output unit analyzes the learner's past responses and outputs audio based on preferred voice tone and speed. For example, the output unit can output optimal audio based on the voice tone the learner has preferred in the past. It can also output optimal audio based on the voice speed to which the learner has responded well in the past. For example, the output unit can output optimal audio based on the voice intonation the learner has preferred in the past. This improves learning effectiveness by providing audio output based on the learner's past responses. Some or all of the above processing in the output unit may be performed using AI, for example, or without AI. For example, the output unit can output audio using an AI model that takes the learner's past responses as input and outputs optimal audio.

[0095] The collection unit can estimate the learner's emotions and adjust the timing of audio collection based on the estimated emotions. For example, the collection unit can estimate emotions from the learner's facial expressions and voice and collect audio at a timing appropriate to those emotions. For example, if the learner is relaxed, the collection unit will collect audio at a natural timing, and if the learner is tense, it will collect audio at a calm timing. It can also collect audio at an appropriate timing if the learner is excited. This improves learning effectiveness by providing audio collection that is tailored to the learner's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the collection unit may be performed using AI, for example, or without AI. For example, the collection unit can collect audio using an AI model that takes the learner's facial expressions and voice data as input, estimates emotions, and collects audio at an appropriate timing.

[0096] The data collection unit can evaluate the accuracy of learners' pronunciation and grammar in real time. For example, the data collection unit can evaluate the accuracy of learners' pronunciation and grammar in real time and provide appropriate feedback. For example, if a learner's pronunciation is inaccurate, the data collection unit can provide real-time feedback, and if their grammar is inaccurate, it can suggest corrections in real time. It can also praise learners in real time if their pronunciation and grammar are accurate. This improves learning effectiveness by evaluating the accuracy of learners' pronunciation and grammar in real time. Some or all of the above processing in the data collection unit may be performed using AI, for example, or without AI. For example, the data collection unit can perform evaluation using an AI model that takes the accuracy of learners' pronunciation and grammar as input, evaluates it in real time, and outputs feedback.

[0097] The audio collection unit can filter out background noise from the learner to collect clear audio. For example, the audio collection unit filters out background noise from the learner's surroundings to remove noise and collect clear audio. For example, if the learner is in a noisy environment, the audio collection unit filters out background noise to collect audio, and if the learner is in a quiet environment, it minimizes background noise to collect audio. It can also filter out wind noise if the learner is outdoors. In this way, clear audio can be collected by filtering out background noise. Some or all of the above processing in the audio collection unit may be performed using AI, for example, or without AI. For example, the audio collection unit can collect audio using an AI model that takes the learner's background noise as input and outputs clear audio by removing noise.

[0098] The collection unit can estimate the learner's emotions and determine the priority of audio to collect based on the estimated emotions. For example, the collection unit can estimate emotions from the learner's facial expressions and voice and determine the priority of audio according to those emotions. For example, if the learner is relaxed, the collection unit will prioritize collecting natural conversational audio, and if the learner is tense, it will prioritize collecting calm audio. It can also prioritize collecting lively audio if the learner is excited. By determining the priority of audio according to the learner's emotions, the learning effect is improved. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the collection unit may be performed using AI, for example, or without AI. For example, the collection unit can collect audio using an AI model that takes learner's facial expressions and voice data as input, estimates emotions, and outputs an appropriate audio priority.

[0099] The collection unit can prioritize collecting audio that is highly relevant, taking into account the learner's geographical location. For example, the collection unit can prioritize collecting audio related to a region, taking into account the learner's geographical location. For instance, if the learner lives in a specific region, the collection unit will prioritize collecting audio related to that region; if the learner is traveling, it will prioritize collecting audio related to their travel destination. Furthermore, if the learner belongs to a specific cultural sphere, it can prioritize collecting audio related to that culture. This improves learning effectiveness by prioritizing the collection of audio based on the learner's geographical location. Some or all of the processing described above in the collection unit may be performed using AI, for example, or without AI. For example, the collection unit can collect audio using an AI model that takes the learner's geographical location as input and outputs highly relevant audio.

[0100] The data collection unit can analyze the learner's social media activity and collect relevant audio. For example, the data collection unit can analyze the learner's social media activity and collect audio related to frequently mentioned topics. For example, the data collection unit can collect audio related to topics that the learner frequently mentions on social media and audio related to the content of accounts that the learner follows. It can also collect audio related to topics in online communities that the learner participates in. This improves learning effectiveness by collecting audio based on the learner's social media activity. Some or all of the processing described above in the data collection unit may be performed using AI, for example, or not. For example, the data collection unit can collect audio using an AI model that takes the learner's social media activity as input and outputs relevant audio.

[0101] The question generation unit can estimate the learner's emotions and adjust the content and tone of the questions based on the estimated emotions. For example, the question generation unit can estimate emotions from the learner's facial expressions and voice and generate questions with content and tone appropriate to those emotions. For example, if the learner is relaxed, the question generation unit will generate questions in a calm tone, and if the learner is excited, it will generate questions in a lively tone. It can also generate questions in a gentle tone if the learner is tired. This improves learning effectiveness by providing questions that are appropriate to the learner's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or a generative AI. The generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the question generation unit may be performed using AI, for example, or without AI. For example, the question generation unit can generate questions using an AI model that takes the learner's facial expressions and voice data as input, estimates emotions, and outputs questions with appropriate content and tone.

[0102] The question generation unit can analyze the learner's past answer history and generate optimal questions. For example, the question generation unit can analyze the learner's past answer history and generate relevant questions. For example, the question generation unit can generate relevant questions based on questions the learner has answered in the past. It can also generate supplementary questions based on questions the learner has struggled with in the past. For example, the question generation unit can generate engaging questions based on topics the learner has shown interest in in the past. This improves learning effectiveness by providing questions based on the learner's past answer history. Some or all of the above processing in the question generation unit may be performed using AI, for example, or without AI. For example, the question generation unit can generate questions using an AI model that takes the learner's past answer history as input and outputs the optimal question.

[0103] The question generation unit can provide customized questions based on the learner's interests and concerns. For example, the question generation unit generates relevant questions based on the learner's interests and concerns. For instance, if the learner is interested in animals, the question generation unit will generate questions about animals; if the learner is interested in sports, it will generate questions related to sports. It can also generate questions about music if the learner is interested in music. This improves learning effectiveness by providing questions based on the learner's interests and concerns. Some or all of the above processing in the question generation unit may be performed using AI, for example, or without AI. For example, the question generation unit can generate questions using an AI model that takes the learner's interests and concerns as input and outputs customized questions.

[0104] The question generation unit can estimate the learner's emotions and adjust the length of the questions based on the estimated emotions. For example, the question generation unit can estimate emotions from the learner's facial expressions and voice and generate questions of appropriate length. For example, the question generation unit can generate short questions when the learner is tired and detailed questions when the learner is focused. It can also generate questions of appropriate length when the learner is excited. This improves learning effectiveness by providing questions of appropriate length according to the learner's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or a generative AI. The generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the question generation unit may be performed using AI or not. For example, the question generation unit can generate questions using an AI model that takes the learner's facial expressions and voice data as input, estimates emotions, and outputs questions of appropriate length.

[0105] The question generation unit can provide highly relevant questions by taking into account the learner's geographical location. For example, the question generation unit can generate questions related to the learner's region by considering their geographical location. For instance, if the learner lives in a specific region, the question generation unit can generate questions related to that region; if the learner is traveling, it can generate questions related to their travel destination. It can also generate questions related to a specific culture if the learner belongs to that culture. This improves learning effectiveness by providing questions based on the learner's geographical location. Some or all of the above processing in the question generation unit may be performed using AI, for example, or without AI. For example, the question generation unit can generate questions using an AI model that takes the learner's geographical location as input and outputs highly relevant questions.

[0106] The question generation unit can analyze the learner's social media activity and include relevant questions. For example, the question generation unit can analyze the learner's social media activity and generate questions related to frequently mentioned topics. For example, the question generation unit can generate questions related to topics that the learner frequently mentions on social media and questions related to the content of accounts that the learner follows. It can also generate questions related to topics in online communities that the learner participates in. This improves learning effectiveness by providing questions based on the learner's social media activity. Some or all of the above processing in the question generation unit may be performed using AI, for example, or not using AI. For example, the question generation unit can generate questions using an AI model that takes the learner's social media activity as input and outputs relevant questions.

[0107] The question output unit can estimate the learner's emotions and adjust the tone and speed of the question's voice based on the estimated emotions. For example, the question output unit can estimate emotions from the learner's facial expressions and voice, and output the question with a tone and speed appropriate to those emotions. For example, if the learner is relaxed, the question output unit will output the question in a relaxed tone, and if the learner is in a hurry, it will output the question at a fast speed. It can also output the question in a bright tone if the learner is excited. This improves learning effectiveness by providing voice output of questions that are appropriate to the learner's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the question output unit may be performed using AI, for example, or without AI. For example, the question output unit can output questions using an AI model that takes the learner's facial expressions and voice data as input, estimates emotions, and outputs questions with an appropriate tone and speed.

[0108] The question output unit can adjust the frequency and volume of the audio based on the learner's auditory characteristics. For example, the question output unit outputs questions at an appropriate frequency and volume, taking into account the learner's auditory characteristics. For instance, if the learner has difficulty hearing high frequencies, the question output unit outputs questions at a low frequency, and if the learner has difficulty hearing low frequencies, it outputs questions at a high frequency. Furthermore, if the learner has a hearing impairment, the question output unit can output questions at an appropriate volume. This improves learning effectiveness by providing audio output of questions tailored to the learner's auditory characteristics. Some or all of the above processing in the question output unit may be performed using AI, for example, or without AI. For example, the question output unit can output questions using an AI model that takes the learner's auditory characteristics as input and outputs questions at an appropriate frequency and volume.

[0109] The question output unit can detect the learner's ambient sounds and output the question at an optimal volume and quality. For example, the question output unit can detect ambient sounds around the learner and output the question at an appropriate volume and quality. For example, if the learner is in a noisy environment, the question output unit will increase the volume of the question and output the question at a lower volume if the learner is in a quiet environment. It can also consider wind noise when the learner is outdoors when outputting the question. This improves learning effectiveness by providing audio output of questions that are tailored to the learner's ambient sounds. Some or all of the above processing in the question output unit may be performed using AI, for example, or without AI. For example, the question output unit can output questions using an AI model that takes the learner's ambient sounds as input and outputs questions at an optimal volume and quality.

[0110] The question output unit can estimate the learner's emotions and adjust the intonation of the question based on the estimated emotions. For example, the question output unit can estimate emotions from the learner's facial expressions and voice and output the question with an intonation appropriate to the emotion. For example, if the learner is relaxed, the question output unit will output the question with a calm intonation, and if the learner is excited, it will output the question with an energetic intonation. It can also output the question with a calm intonation if the learner is tired. This improves learning effectiveness by providing question intonation that matches the learner's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is a text generation AI (e.g., LLM) or a multimodal generation AI, but is not limited to these examples. Some or all of the above processing in the question output unit may be performed using AI, for example, or without AI. For example, the question output unit can output questions using an AI model that takes the learner's facial expressions and voice data as input, estimates emotions, and outputs questions with appropriate intonation.

[0111] The question output unit can select the optimal audio format considering the learner's device information. For example, the question output unit can select an appropriate audio format based on the type of device the learner is using. For instance, if the learner is using a smartphone, the question output unit will select the audio format best suited for smartphones; if they are using a tablet, it will select the audio format best suited for tablets. It can also select the audio format best suited for PCs if the learner is using a PC. This improves learning effectiveness by providing audio output of questions tailored to the learner's device information. Some or all of the above processing in the question output unit may be performed using AI, for example, or without AI. For example, the question output unit can output questions using an AI model that takes the learner's device information as input and outputs questions in the optimal audio format.

[0112] The question output unit can analyze the learner's past responses and output the optimal voice. For example, the question output unit analyzes the learner's past responses and outputs questions based on preferred voice tone and speed. For example, the question output unit outputs questions in the optimal voice based on the voice tone the learner has preferred in the past. It can also output questions in the optimal voice based on the voice speed the learner has responded well to in the past. For example, the question output unit outputs questions in the optimal voice based on the voice intonation the learner has preferred in the past. This improves learning effectiveness by providing voice output of questions based on the learner's past responses. Some or all of the above processing in the question output unit may be performed using AI, for example, or without AI. For example, the question output unit can output questions using an AI model that takes the learner's past responses as input and outputs questions in the optimal voice.

[0113] The language selection unit can estimate the learner's emotions and suggest language choices based on those emotions. For example, the language selection unit can estimate emotions from the learner's facial expressions and voice and suggest language choices that correspond to those emotions. For instance, if the learner is relaxed, the language selection unit suggests a language that indicates high learning motivation; if the learner is excited, it suggests a language that will pique their interest. It can also suggest a simple language if the learner is tired. This improves learning effectiveness by providing language choices that correspond to the learner's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or a generative AI. The generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processing in the language selection unit may be performed using AI, or not. For example, the language selection unit can suggest language choices using an AI model that takes the learner's facial expressions and voice data as input, estimates emotions, and outputs appropriate language choices.

[0114] The language selection unit can analyze the learner's past learning history and suggest the most suitable language. For example, the language selection unit can analyze the learner's past learning history and suggest relevant languages. For example, the language selection unit can suggest relevant languages ​​based on the languages ​​the learner has learned in the past. It can also suggest supplementary languages ​​based on languages ​​the learner has struggled with in the past. For example, the language selection unit can suggest languages ​​that will interest the learner based on languages ​​the learner has shown interest in in the past. This improves learning effectiveness by providing language selection suggestions based on the learner's past learning history. Some or all of the above processing in the language selection unit may be performed using AI, for example, or without AI. For example, the language selection unit can make language selection suggestions using an AI model that takes the learner's past learning history as input and outputs the most suitable language.

[0115] The language selection unit can estimate the learner's emotions and determine language selection priorities based on the estimated emotions. For example, the language selection unit can estimate emotions from the learner's facial expressions and voice and determine language selection priorities according to those emotions. For example, if the learner is relaxed, the language selection unit will prioritize suggesting languages ​​that indicate a high level of learning motivation, and if the learner is excited, it will prioritize suggesting languages ​​that will pique their interest. It can also prioritize suggesting simpler languages ​​if the learner is tired. This improves learning effectiveness by providing language selection priorities that correspond to the learner's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the language selection unit may be performed using AI, or not using AI. For example, the language selection unit can determine language selection priorities using an AI model that takes the learner's facial expressions and voice data as input, estimates emotions, and outputs appropriate language selection priorities.

[0116] The language selection unit can suggest highly relevant languages ​​by considering the learner's geographical location. For example, the language selection unit can suggest the language used in a particular region, taking into account the learner's geographical location. For example, if the learner lives in a specific region, the language selection unit can suggest the language used in that region; if the learner is traveling, it can suggest the language used in their travel destination. It can also suggest the language used in a particular culture if the learner belongs to that culture. This improves learning effectiveness by providing language selection suggestions based on the learner's geographical location. Some or all of the above processing in the language selection unit may be performed using AI, for example, or without AI. For example, the language selection unit can suggest language selections using an AI model that takes the learner's geographical location as input and outputs highly relevant languages.

[0117] The assessment unit can estimate the learner's emotions and adjust the learning progress criteria based on the estimated emotions. For example, the assessment unit can estimate emotions from the learner's facial expressions and voice and set criteria according to those emotions. For example, the assessment unit can apply strict criteria when the learner is relaxed and lenient criteria when the learner is excited. It can also apply gentle criteria when the learner is tired. This improves learning effectiveness by providing learning progress criteria that are appropriate to the learner's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the assessment unit may be performed using AI, or not using AI. For example, the assessment unit can adjust the learning progress criteria using an AI model that takes the learner's facial expressions and voice data as input, estimates emotions, and outputs appropriate criteria.

[0118] The evaluation unit can analyze the learner's past learning history and set optimal evaluation criteria. For example, the evaluation unit can analyze the learner's past learning history and set appropriate evaluation criteria. For example, the evaluation unit can set appropriate evaluation criteria based on what the learner has learned in the past. It can also set supplementary evaluation criteria based on what the learner has struggled with in the past. For example, the evaluation unit can set attractive evaluation criteria based on what the learner has shown interest in in the past. This improves learning effectiveness by providing evaluation criteria based on the learner's past learning history. Some or all of the above processing in the evaluation unit may be performed using AI, for example, or without AI. For example, the evaluation unit can set evaluation criteria using an AI model that takes the learner's past learning history as input and outputs optimal evaluation criteria.

[0119] The assessment unit can estimate the learner's emotions and adjust the order in which the learning progress assessment results are displayed based on the estimated emotions. For example, the assessment unit can estimate emotions from the learner's facial expressions and voice and display the assessment results in an order appropriate to the emotions. For example, if the learner is relaxed, the assessment unit can display detailed results first, and if the learner is excited, it can display concise results first. It can also display brief results first if the learner is tired. This improves learning effectiveness by providing learning progress assessment results that are appropriate to the learner's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the assessment unit may be performed using AI, for example, or without AI. For example, the assessment unit can adjust the order in which the assessment results are displayed using an AI model that takes the learner's facial expressions and voice data as input, estimates emotions, and displays the assessment results in an appropriate order.

[0120] The evaluation unit can set highly relevant evaluation criteria by considering the learner's geographical location information. For example, the evaluation unit can set evaluation criteria used in a particular region by considering the learner's geographical location information. For example, if the learner lives in a particular region, the evaluation unit can set evaluation criteria used in that region, and if the learner is traveling, it can set evaluation criteria used in the travel destination. It can also set evaluation criteria used in a particular culture if the learner belongs to a particular cultural area. This improves learning effectiveness by providing evaluation criteria based on the learner's geographical location information. Some or all of the above processing in the evaluation unit may be performed using AI, for example, or without using AI. For example, the evaluation unit can set evaluation criteria using an AI model that takes the learner's geographical location information as input and outputs highly relevant evaluation criteria.

[0121] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0122] Language learning programs can provide customized lectures based on the learner's learning style. For example, they can generate lectures that heavily utilize images and videos for visual learners, and lectures that heavily utilize audio and music for auditory learners. They can also generate lectures that include interactive elements for tactile learners. By providing lectures tailored to the learner's learning style, learning effectiveness is improved. The lecture generation unit can generate lectures using an AI model that takes the learner's learning style as input and outputs the most suitable lecture content.

[0123] Language learning programs can provide real-time feedback based on the learner's progress. For example, if a learner completes a specific task, they can immediately receive positive feedback; if they are struggling with a task, they can receive additional hints and support. Furthermore, once a learner achieves a certain level of progress, guidance can be provided to help them move on to the next step. This improves learning effectiveness by providing feedback tailored to the learner's progress. The feedback system uses an AI model that takes the learner's progress as input and outputs appropriate feedback.

[0124] Language learning programs can provide customized learning plans based on the learner's learning objectives. For example, if a learner aims to pass a specific exam, a learning plan tailored to that exam can be provided; if their goal is to acquire conversational skills, a learning plan focusing on practical conversation practice can be provided. Furthermore, if a learner wants to learn business English, a learning plan tailored to business situations can be provided. This improves learning effectiveness by providing learning plans that match the learner's objectives. The learning plan generation unit can generate learning plans using an AI model that takes the learner's learning objectives as input and outputs the optimal learning plan.

[0125] The language learning program can suggest the optimal learning method based on the learner's learning environment. For example, if a learner is studying at home, it can suggest a learning method that takes place in a quiet environment; if a learner is studying during their commute, it can suggest a method that allows them to study while on the go. Furthermore, if a learner is studying in a group, it can suggest a learning method that incorporates group discussions and collaborative work. This improves learning effectiveness by providing learning methods tailored to the learner's environment. The learning method suggestion unit can suggest learning methods using an AI model that takes the learner's learning environment as input and outputs the optimal learning method.

[0126] The language learning program can suggest review timings based on the learner's learning history. For example, if a certain period has passed since the learner studied a particular topic, it can suggest a review of that topic to reinforce what they have learned. It can also suggest regular reviews of topics that the learner has struggled with in the past. By providing review timings based on the learner's learning history, the learning effect is improved. The review suggestion unit can suggest review timings using an AI model that takes the learner's learning history as input and outputs the optimal review timing.

[0127] Language learning programs can estimate the learner's emotions and adjust the learning content based on those emotions. For example, if a learner is stressed, relaxing content can be provided; if a learner is excited, challenging content can be provided. Furthermore, if a learner is tired, content that can be completed in a short time can be provided. This improves learning effectiveness by providing learning content tailored to the learner's emotions. The emotion estimation unit uses an AI model that takes the learner's facial expressions and voice data as input, estimates emotions, and outputs appropriate learning content to adjust the learning content.

[0128] The language learning program can estimate the learner's emotions and adjust the learning pace based on those emotions. For example, if the learner is relaxed, the learning progresses at a normal pace; if the learner is excited, it progresses at a faster pace. Furthermore, if the learner is tired, the learning can be slowed down. This improves learning effectiveness by providing a learning pace that matches the learner's emotions. The emotion estimation unit uses an AI model that takes the learner's facial expressions and voice data as input, estimates their emotions, and outputs an appropriate learning pace to adjust the learning speed.

[0129] The language learning program can estimate the learner's emotions and suggest break times based on those emotions. For example, if the learner is tired, it can suggest a break and provide time to refresh. Conversely, if the learner is focused, it can suggest continuing the learning process. By providing break times that match the learner's emotions, the learning effect is improved. The emotion estimation unit can suggest break times using an AI model that takes the learner's facial expressions and voice data as input, estimates their emotions, and outputs appropriate break times.

[0130] Language learning programs can estimate learners' emotions and provide support to maintain their motivation based on those estimated emotions. For example, if a learner is losing motivation, it can provide encouraging messages; if a learner is highly motivated, it can suggest further challenges. It can also provide relaxing content if a learner is feeling stressed. By providing motivational support tailored to the learner's emotions, learning effectiveness can be improved. The emotion estimation unit uses an AI model that takes the learner's facial expressions and voice data as input, estimates their emotions, and outputs appropriate motivational support to provide support for maintaining motivation.

[0131] Language learning programs can estimate the learner's emotions and adjust learning feedback based on those emotions. For example, if the learner is relaxed, detailed feedback is provided; if the learner is excited, concise feedback is provided. Furthermore, if the learner is tired, feedback can be provided in a gentle tone. This improves learning effectiveness by providing feedback tailored to the learner's emotions. The emotion estimation unit uses an AI model that takes the learner's facial expressions and voice data as input, estimates emotions, and outputs appropriate feedback to adjust the feedback accordingly.

[0132] The following briefly describes the processing flow for example form 2.

[0133] Step 1: The lecture generation unit generates audio for lectures on foreign languages. For example, it uses a generation AI to generate lectures on foreign language grammar and pronunciation. The generation AI can generate lecture content using a text generation AI (e.g., LLM). It can also use the generation AI to generate lectures on foreign language conversation practice. For example, the generation AI generates a conversation scenario and delivers a lecture based on it. Step 2: The output unit outputs the audio generated by the lecture generation unit through a speaker attached to the stuffed animal. For example, audio generated using a generation AI can be output through the speaker. The output unit can also use the generation AI to adjust the tone and speed of the audio before outputting it. For example, the tone and speed of the audio can be adjusted according to the child's age and interests. Step 3: The collection unit collects the participants' voices. For example, it can collect participants' voices using a microphone. It can also analyze the collected audio using generative AI. For example, it can analyze the accuracy of the participants' pronunciation and grammar. Step 4: The question generation unit generates appropriate questions based on the audio collected by the collection unit. For example, using a generation AI, questions can be generated based on the learner's pronunciation and grammatical accuracy. Alternatively, the generation AI can be used to generate questions based on the learner's age and interests. For example, if a child is interested in animals, the AI ​​can generate a question such as, "What is your favorite animal?" Step 5: The question output unit outputs the audio of the question generated by the question generation unit through the speaker. For example, the audio of a question generated using a generation AI can be output through the speaker. It is also possible to adjust the tone and speed of the audio using the generation AI. For example, the tone and speed of the audio can be adjusted according to the child's age and interests.

[0134] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0135] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include text generation AI, image generation AI, and multimodal generation AI. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats from audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each of the above parts is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example.Furthermore, processing performed by AI, including generative AI, may be replaced with rule-based processing, and rule-based processing may be replaced with processing performed by AI, including generative AI.

[0136] Furthermore, the processing performed by the data processing system 10 described above is carried out by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it may also be carried out by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0137] For example, the lecture generation unit is implemented by the specific processing unit 290 of the data processing device 12. For example, the output unit is implemented by the control unit 46A of the smart device 14. For example, the collection unit is implemented by the microphone 38B of the smart device 14. For example, the question generation unit is implemented by the specific processing unit 290 of the data processing device 12. For example, the question output unit is implemented by the control unit 46A of the smart device 14. The correspondence between each unit and the devices and control units is not limited to the examples described above, and various modifications are possible.

[0138] [Second Embodiment] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0139] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0140] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0141] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0142] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0143] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0144] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0145] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing by the processor 28. The storage 32 stores the specific processing program 56.

[0146] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0147] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0148] In the smart glasses 214, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0149] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0150] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0151] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0152] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart glasses 214 or an external device, and the smart glasses 214 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0153] For example, the lecture generation unit is implemented by the specific processing unit 290 of the data processing device 12. For example, the output unit is implemented by the control unit 46A of the smart glasses 214. For example, the data collection unit is implemented by the microphone 238 of the smart glasses 214. For example, the question generation unit is implemented by the specific processing unit 290 of the data processing device 12. For example, the question output unit is implemented by the control unit 46A of the smart glasses 214. The correspondence between each unit and the device or control unit is not limited to the examples described above, and various modifications are possible.

[0154] [Third Embodiment] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0155] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0156] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0157] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0158] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0159] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0160] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0161] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0162] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0163] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0164] In the headset terminal 314, specific processing is performed by the processor 46. The storage 50 stores a specific program 60. The processor 46 reads the specific program 60 from the storage 50 and executes the read specific program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific program 60 executed on the RAM 48. The headset terminal 314 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0165] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0166] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0167] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0168] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset terminal 314, but may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset terminal 314. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the headset terminal 314 or an external device, and the headset terminal 314 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0169] For example, the lecture generation unit is implemented by the specific processing unit 290 of the data processing device 12. For example, the output unit is implemented by the control unit 46A of the headset terminal 314. For example, the collection unit is implemented by the microphone 238 of the headset terminal 314. For example, the question generation unit is implemented by the specific processing unit 290 of the data processing device 12. For example, the question output unit is implemented by the control unit 46A of the headset terminal 314. The correspondence between each unit and the device or control unit is not limited to the examples described above, and various modifications are possible.

[0170] [Fourth Embodiment] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0171] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0172] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0173] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0174] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0175] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS image sensor or CCD image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0176] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0177] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. The robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0178] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0179] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0180] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0181] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in storage 50. Processor 46 reads the specific program 60 from storage 50 and executes it on RAM 48. The specific processing is achieved by processor 46 acting as a control unit 46A according to the specific program 60 executed on RAM 48. Robot 414 also has data generation model 58 and emotion identification model 59, similar to those of the robot, and can perform processing similar to that of the specific processing unit 290 using these models.

[0182] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0183] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0184] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0185] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the robot 414 or an external device, and the robot 414 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0186] For example, the lecture generation unit is implemented by the specific processing unit 290 of the data processing device 12. For example, the output unit is implemented by the control unit 46A of the stuffed animal 414. For example, the collection unit is implemented by the microphone 238 of the stuffed animal 414. For example, the question generation unit is implemented by the specific processing unit 290 of the data processing device 12. For example, the question output unit is implemented by the control unit 46A of the stuffed animal 414. The correspondence between each unit and the device or control unit is not limited to the examples described above, and various modifications are possible.

[0187] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0188] Figure 9 shows the emotion map 400, in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0189] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0190] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0191] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, and motorcycles, emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0192] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0193] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0194] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes computer 22 and multiple other computers.

[0195] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0196] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0197] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0198] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0199] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0200] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0201] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0202] Furthermore, although the above-described examples were divided into four embodiments, some or all of these embodiments may be combined. Also, the smart device 14, smart glasses 214, headset terminal 314, and robot 414 are just examples, and they may be combined, or other devices may be used. Also, although the above-described examples were divided into two embodiments, Embodiment 1 and Embodiment 2, these may be combined.

[0203] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and other things that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0204] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0205] (Note 1) A lecture generation unit that generates audio for lectures on foreign languages, An output unit that outputs the audio generated by the lecture generation unit from a speaker attached to the stuffed animal, A collection department that gathers feedback from participants, A question generation unit generates specific questions based on the audio collected by the aforementioned collection unit, The system includes a question output unit that outputs the audio of the question generated by the question generation unit from the speaker. A system characterized by the following features. (Note 2) The aforementioned question generation unit, Questions are generated based on the participant's age, specific interests, and past conversation history. The system described in Appendix 1, characterized by the features described herein. (Note 3) The aforementioned question generation unit, The system evaluates the learner's pronunciation and grammar, and generates questions based on the specific evaluation results. The system described in Appendix 1, characterized by the features described herein. (Note 4) The aforementioned question output unit is: It includes a language selection section to support specific languages. The system described in Appendix 1, characterized by the features described herein. (Note 5) The aforementioned question generation unit, It estimates the specific emotions of the participants and generates questions based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 6) The system further includes a determination unit that analyzes the audio of the answer to the question output by the question output unit and determines the learner's specific level of learning based on the analysis results. The aforementioned lecture generation unit, The difficulty level of foreign language lectures will be adjusted based on the specific learning progress of the students. The system described in Appendix 1, characterized by the features described herein. (Note 7) The aforementioned lecture generation unit, The system estimates the specific emotions of the participants and adjusts the content and tone of the lecture based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 8) The aforementioned lecture generation unit, Analyze the student's past learning history to generate specific lecture content. The system described in Appendix 1, characterized by the features described herein. (Note 9) The aforementioned lecture generation unit, We provide customized lectures based on the specific interests and concerns of the participants. The system described in Appendix 1, characterized by the features described herein. (Note 10) The aforementioned lecture generation unit, The system estimates the specific emotions of the participants and adjusts the length of the lecture based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 11) The aforementioned lecture generation unit, Provide specific content that takes into account the geographical location of the participants. The system described in Appendix 1, characterized by the features described herein. (Note 12) The aforementioned lecture generation unit, Analyze participants' social media activity and include specific topics. The system described in Appendix 1, characterized by the features described herein. (Note 13) The output unit is, It estimates the specific emotions of the participants and adjusts the tone and speed of the voice based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 14) The output unit is, The audio frequency and volume are adjusted based on the specific auditory characteristics of each participant. The system described in Appendix 1, characterized by the features described herein. (Note 15) The output unit is, The system detects ambient sounds from the participant's environment and outputs specific audio. The system described in Appendix 1, characterized by the features described herein. (Note 16) The output unit is, It estimates the specific emotions of the participants and adjusts the intonation of their voices based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 17) The output unit is, Select a specific audio format considering the participant's device information. The system described in Appendix 1, characterized by the features described herein. (Note 18) The output unit is, Analyze past responses from participants and produce specific audio output. The system described in Appendix 1, characterized by the features described herein. (Note 19) The aforementioned collection unit is We estimate the specific emotions of the participants and adjust the timing of audio collection based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 20) The aforementioned collection unit is The system evaluates the specific accuracy of students' pronunciation and grammar in real time. The system described in Appendix 1, characterized by the features described herein. (Note 21) The aforementioned collection unit is Filter the background noise of the participants to collect specific audio. The system described in Appendix 1, characterized by the features described herein. (Note 22) The aforementioned collection unit is The system estimates the specific emotions of the participants and determines the priority of audio recordings to be collected based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 23) The aforementioned collection unit is Prioritize collecting specific audio recordings, taking into account the geographical location of the participants. The system described in Appendix 1, characterized by the features described herein. (Note 24) The aforementioned collection unit is Analyze participants' social media activity and collect specific audio recordings. The system described in Appendix 1, characterized by the features described herein. (Note 25) The aforementioned question generation unit, We estimate the specific emotions of the participants and adjust the content and tone of the questions based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 26) The aforementioned question generation unit, Analyze the participants' past response history to generate specific questions. The system described in Appendix 1, characterized by the features described herein. (Note 27) The aforementioned question generation unit, Provide customized questions based on the participants' specific interests and concerns. The system described in Appendix 1, characterized by the features described herein. (Note 28) The aforementioned question generation unit, The system estimates the specific emotions of the participants and adjusts the length of the questions based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 29) The aforementioned question generation unit, Provide specific questions that take into account the geographical location of the participants. The system described in Appendix 1, characterized by the features described herein. (Note 30) The aforementioned question generation unit, Analyze participants' social media activity and include specific questions. The system described in Appendix 1, characterized by the features described herein. (Note 31) The aforementioned question output unit is: The system estimates the specific emotions of the participants and adjusts the tone and speed of the questions based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 32) The aforementioned question output unit is: The audio frequency and volume are adjusted based on the specific auditory characteristics of each participant. The system described in Appendix 1, characterized by the features described herein. (Note 33) The aforementioned question output unit is: The system detects ambient sounds from the participant's environment and outputs specific audio. The system described in Appendix 1, characterized by the features described herein. (Note 34) The aforementioned question output unit is: Estimate the specific emotions of the participants and adjust the intonation of the questions based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 35) The aforementioned question output unit is: Select a specific audio format considering the participant's device information. The system described in Appendix 1, characterized by the features described herein. (Note 36) The aforementioned question output unit is: Analyze past responses from participants and produce specific audio output. The system described in Appendix 1, characterized by the features described herein. (Note 37) The language selection unit, The system estimates the specific emotions of the participants and suggests language choices based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 38) The language selection unit, We analyze the student's past learning history and suggest specific languages. The system described in Appendix 1, characterized by the features described herein. (Note 39) The language selection unit, The system estimates the specific emotions of the participants and determines the priority of language selection based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 40) The language selection unit, We will propose a specific language based on the geographical location of the participants. The system described in Appendix 1, characterized by the features described herein. (Note 41) The determination unit, The system estimates the specific emotions of the participants and adjusts the learning progress criteria based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 42) The determination unit, Analyze the past learning history of participants and set specific evaluation criteria. The system described in Appendix 1, characterized by the features described herein. (Note 43) The determination unit, The system estimates the specific emotions of the participants and adjusts the order in which the learning progress assessment results are displayed based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 44) The determination unit, Specific evaluation criteria are set considering the geographical location information of the participants. The system described in Appendix 1, characterized by the features described herein. [Explanation of symbols]

[0206] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots

Claims

1. A lecture generation unit that generates audio for lectures on foreign languages, An output unit that outputs the audio generated by the lecture generation unit from a speaker attached to the stuffed animal, A collection department that gathers feedback from participants, A question generation unit generates specific questions based on the audio collected by the aforementioned collection unit, The system includes a question output unit that outputs the audio of the question generated by the question generation unit from the speaker. A system characterized by the following features.

2. The aforementioned question generation unit, Questions are generated based on the participant's age, specific interests, and past conversation history. The system according to feature 1.

3. The aforementioned question generation unit, The system evaluates the learner's pronunciation and grammar, and generates questions based on the specific evaluation results. The system according to feature 1.

4. The aforementioned question output unit is: It includes a language selection section to support specific languages. The system according to feature 1.

5. The aforementioned question generation unit, It estimates the specific emotions of the participants and generates questions based on those estimated emotions. The system according to feature 1.

6. The system further includes a determination unit that analyzes the audio of the answer to the question output by the question output unit and determines the learner's specific level of learning based on the analysis results. The aforementioned lecture generation unit, The difficulty level of foreign language lectures will be adjusted based on the specific learning progress of the students. The system according to feature 1.

7. The aforementioned lecture generation unit, The system estimates the specific emotions of the participants and adjusts the content and tone of the lecture based on those estimated emotions. The system according to feature 1.

8. The aforementioned lecture generation unit, Analyze the student's past learning history to generate specific lecture content. The system according to feature 1.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A