system

The system addresses teacher workload and educational quality by automating audio lessons and student question responses using digital educational materials, improving lesson management and student interaction.

JP2026073378APending Publication Date: 2026-05-01SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-18
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Teachers face significant burdens in lesson preparation, progress management, and responding to student questions, leading to increased overtime and reduced educational quality.

Method used

A system that automatically generates and conducts audio lessons using digital educational materials, includes input receiving and response generation means to reduce teacher workload and improve educational quality.

Benefits of technology

The system efficiently manages lesson progression and provides immediate responses to student questions, reducing teacher burden and enhancing educational quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026073378000001_ABST
    Figure 2026073378000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] Information storage means for providing digital educational materials, A generation means that receives digital educational materials from the aforementioned information storage means and generates audio, Audio output means for outputting audio generated by the generation means, An input receiving means that receives student voice input and generates question data, A response generation means that generates an audio response corresponding to the aforementioned question data, A response output means for reproducing the aforementioned voice response, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Currently, in the educational field, the burden on teachers is extremely large. In particular, a great deal of time and effort are spent on lesson preparation, progress, and responses to questions. As a result, there is a social problem that teachers' overtime work hours are increasing, and their mental and physical burdens are growing. This is making it increasingly difficult to maintain the quality of education. The present invention aims to solve these problems.

Means for Solving the Problems

[0005] This invention provides a system that automatically generates and conducts audio lessons using digital educational materials. The system includes an information storage means for storing digital educational materials, a generation means for generating audio necessary for lessons, and an audio output means for outputting the generated audio. It also includes an input receiving means for receiving questions from students, a response generation means for generating responses based on those questions, and a response output means for outputting the generated responses. With this configuration, it is possible to reduce the burden on teachers in conducting lessons and answering questions, thereby improving the quality of education while reducing the workload of teachers.

[0006] "Digital educational materials" refer to educational information such as teaching materials and textbooks that are recorded in digital data format.

[0007] "Information storage means" refers to devices and systems that can store digital educational materials and provide them as needed.

[0008] "Generation means" refers to devices or algorithms that have the function of creating audio lessons based on digital educational materials.

[0009] "Audio output means" refers to a device that has the function of playing back generated audio data as audio that can be heard in the real world.

[0010] An "input receiving means" refers to a device or system that has the function of receiving voice input from students and processing it as digital data.

[0011] A "response generation means" refers to a device or algorithm that has the function of generating appropriate answers from digital educational materials based on student question data.

[0012] A "response output means" is a device that has the function of outputting the generated response as audio. [Brief explanation of the drawing]

[0013] [Figure 1]This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]

[0014] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0015] First, the terms used in the following description will be explained.

[0016] In the following embodiments, the labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0017] In the following embodiments, the labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0018] In the following embodiments, the labeled storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0019] In the following embodiments, the labeled communication I / F (Interface) is an interface including a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), etc.

[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0021] [First Embodiment]

[0022] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0023] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0024] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0025] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0026] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0028] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0029] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0030] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0031] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0032] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0033] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0034] This invention is a system for managing lesson progress and responding to student questions based on digitized educational materials. The following describes a specific implementation of this system.

[0035] First, at the heart of the system resides digital educational materials, and the information storage means that stores these materials acts as a server. The server accesses the digital educational materials and, based on this, supplies the audio necessary for the lesson to the audio generation means. The server provides the necessary educational content according to the data requested by the audio generation means and distributes the generated audio data to the terminals.

[0036] Next, the terminal plays the audio data received from the server and conducts the classroom lesson through the audio output device. The terminal selects appropriate audio according to the context of the lesson and supports the flow of the lesson. The terminal is also equipped with an input device to receive questions from students via voice. The questions entered via voice are converted into digital data and sent to the server.

[0037] The server generates appropriate answers from digital educational materials based on the received question data. To do this, the server uses response generation means and AI algorithms to quickly construct appropriate answers. The answers are then sent back to the terminal as audio data.

[0038] The terminal provides responses to student questions by playing back audio data received from the server using its response output mechanism. This allows for immediate feedback on student inquiries without interrupting the flow of the lesson.

[0039] The user (teacher) can monitor the progress of the lesson, intervene in the device as needed, and pause or add explanations to the audio. This allows the teacher to confirm students' understanding and take on roles such as supplementing key points.

[0040] As a concrete example, when lecturing on "Edo period culture" in a history class, the server generates audio for the relevant topic from a digitized history textbook, and the terminals play this audio in the classroom. If a student asks, "I want to know more about Edo period clothing," the terminal transmits this as voice input to the server, which uses a response generation mechanism to provide a detailed explanation and returns it to the student as voice output.

[0041] Thus, this system is designed to efficiently support lessons while reducing the burden on teachers.

[0042] The following describes the processing flow.

[0043] Step 1:

[0044] The server stores digital educational materials in an information storage device and makes them accessible as needed. The server organizes teaching material data related to topics used in lessons and prepares it so that the generation device can efficiently generate audio.

[0045] Step 2:

[0046] At the start of a class, the user operates their device and selects a class topic. This action causes the device to send a request to the server based on the selected topic and begin retrieving the necessary educational materials.

[0047] Step 3:

[0048] The server, in response to the user's request, passes educational materials on the selected topic to the audio generation device. The server then converts the specified content into audio data and prepares it for transmission to the terminal.

[0049] Step 4:

[0050] The terminal plays back audio data received from the server using an audio output device. This allows the lesson to begin with an audio guide such as, "Today we will learn about XX."

[0051] Step 5:

[0052] When a student asks a question, the terminal's input receiving mechanism receives it as voice input and converts it into digital data. The terminal then sends the converted question data to the server and requests the generation of a response.

[0053] Step 6:

[0054] The server receives the question data and uses a response generation mechanism to create an answer based on digital educational materials. The server generates the optimal answer to the question and sends it to the terminal as audio data.

[0055] Step 7:

[0056] The device plays back the received audio data of the response via a response output device, providing the student with the answer to their question. This allows students to resolve their questions without disrupting the flow of the lesson.

[0057] Step 8:

[0058] The user monitors the lesson progress and controls audio playback via their device as needed. The user can provide supplementary explanations and adjust the lesson pace based on student reactions.

[0059] Step 9:

[0060] When the lesson ends, the user exits the lesson by operating their terminal. The server creates a lesson log and stores data to help improve future lessons.

[0061] (Example 1)

[0062] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0063] In today's educational environment, teachers spend a great deal of time and effort managing the progress of lessons and responding to individual student questions. As a result, there is a risk of insufficient time for home education and a decline in the quality of lesson content. In particular, the burden on teachers increases when individualized support tailored to each student's level of understanding is required. To address these challenges, a system is needed that allows for efficient lesson progression and immediate responses to individual student questions.

[0064] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0065] In this invention, the server includes a storage device means for providing digital teaching materials, a speech synthesis device means for acquiring digital teaching materials from the storage device and generating audio, and an input device means for receiving voice input from learners and generating inquiry data. This makes it possible to respond to students' questions in real time while smoothly conducting lessons.

[0066] A "storage device" is a device used to store and keep digital learning materials accessible.

[0067] A "speech synthesis device" is a device that generates audio data from information acquired from digital educational materials.

[0068] A "playback device" is a device that outputs generated audio data and allows learners to listen to it.

[0069] An "input device" is a device that receives voice input from learners and processes it as inquiry data.

[0070] A "response generation device" is a device that generates an appropriate voice response based on inquiry data.

[0071] An "output device" is a device that plays back the generated voice response and provides it to the learner.

[0072] A "generative artificial intelligence model" is a computational model that utilizes information technology to generate optimized responses based on query data.

[0073] This invention relates to an educational support system that utilizes digital teaching materials, and is configured in which a server, a terminal, and a user work together in a coordinated manner.

[0074] The server is the core of the educational support system. It stores digital learning materials in cloud storage and functions as a memory device. When a request for learning materials is received, the server uses a speech synthesis device and utilizes speech technologies such as Google® Text-to-Speech API to convert the materials into audio data. The converted audio data is then streamed to the terminal.

[0075] The terminal is a device that plays audio data received from the server. Equipped with a playback device, the terminal supports the progress of lessons by outputting audio within the classroom. The terminal also has an input device that receives audio input from learners and transmits it to the server. The input questions are processed appropriately and converted into digital data.

[0076] The user (teacher), acting as the system's supervisor, pauses audio output and provides additional explanations via their terminal as needed during lessons. This allows for flexible lesson progression based on the learners' understanding. The user also reviews learners' questions and provides supplementary information where further explanation is required.

[0077] As a concrete example, consider a scenario in a history class where a lecture is given on "Medieval European Culture." The server generates audio for the relevant topic from digitized history materials, and the terminals play this audio in the classroom. For example, if a student asks, "Please tell me about medieval European clothing," the terminal uses speech recognition technology to transmit the question to the server. The server uses a response generation device and a generative AI model to generate an appropriate answer. This answer is then converted back into audio data and provided to the student via the terminal.

[0078] As a concrete example of a prompt, by inputting "Please describe medieval European clothing in detail," the generating AI model can extract relevant information and provide a detailed explanation. In this way, the present invention realizes improved efficiency and interactive learning support in educational settings.

[0079] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0080] Step 1:

[0081] The server retrieves digital learning materials from cloud storage and saves them to its storage device. The input includes the file format of the digital learning materials. Upon receiving a request to retrieve materials, the server searches for the relevant content and prepares the data. The output is a dataset in which the learning material content is available for access.

[0082] Step 2:

[0083] The server processes the educational material data using a text-to-speech (Speech Synthesizer) to convert it into audio data. Specifically, it utilizes speech technologies such as the Google Text-to-Speech API to convert text data into audio files. In this process, the text data is input into the Speech Synthesizer and output as audio data.

[0084] Step 3:

[0085] The server streams the generated audio data to the terminals. This ensures that audio for use in class is delivered in real time. The input is the converted audio data, and the output is the audio streaming to the terminals via the network.

[0086] Step 4:

[0087] The terminal receives audio data from the server and outputs it within the classroom via a playback device. For playback to begin, audio data must be input to the terminal. Specifically, the audio is played throughout the classroom via speakers. The output is the audio audible to the class participants.

[0088] Step 5:

[0089] The terminal receives voice input from the learner and converts it into query data using an input device. The input includes the learner's question, which is collected via a microphone and processed into digital data. The resulting output is query data that is sent to the server.

[0090] Step 6:

[0091] The server processes the query data using a response generator and generates an appropriate response using a generation AI model. This process involves designing a prompt and inputting it into the model to obtain a response in appropriate text format. The output is the response information that should be converted into audio data.

[0092] Step 7:

[0093] The server converts the generated text responses into speech and then sends them to the terminal. A text-to-speech synthesis device is used for this conversion. The input is the response generated by the AI ​​model, and the output is the audio data streamed to the terminal.

[0094] Step 8:

[0095] The terminal finally outputs the audio data received from the server back into the classroom via a playback device. This completes the feedback process to the learners. The input is audio data, and the output is an audible response.

[0096] (Application Example 1)

[0097] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0098] Current education systems struggle to provide interactive learning experiences in the classroom, particularly in real-time question-and-answer sessions where quick and accurate feedback is difficult to obtain. Furthermore, the audio guides and explanations users receive when searching for information in virtual environments are limited, leaving much room for improvement in the depth of learning. This, in turn, limits the provision of efficient learning environments and the quality of education, posing significant challenges.

[0099] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0100] In this invention, the server includes data storage means, voice generation means, and voice output means. This enables users to search for information in a virtual environment in real time and receive immediate responses from AI.

[0101] A "data storage device" is a device that can store digital learning materials and retrieve them as needed.

[0102] A "speech generation means" is a device that creates audio data based on stored digital learning materials.

[0103] "Audio output means" refers to a device for physically playing back the generated audio data and providing it to the user.

[0104] A "voice reception device" is a device that recognizes voice input from a user and converts it into digital data.

[0105] A "response generation device" is a device that constructs and responds to appropriate information in real time based on inquiries from users.

[0106] A "response output device" is a device for reproducing the generated response and communicating it to the user.

[0107] An "exploration tool" is a device that provides guidance to users when they search for and learn information in a virtual space.

[0108] An "interactive response system" is a device that uses AI to enable immediate responses to user questions.

[0109] To implement this invention, a system is configured to support learning and exploration in a virtual environment. The server stores digital learning materials and uses a speech generation means to create audio data related to the information requested by the user. The speech output means transmits this audio data to a terminal and plays it back in a format audible to the user. The terminal is equipped with a speech receiving means that receives the user's voice input, converts it into digital data, and transmits it to the server.

[0110] The server utilizes response generation mechanisms and constructs instant responses using AI algorithms. The generated responses are then converted back into audio data and delivered to the user through response output mechanisms. This allows users to gain an interactive learning experience in a virtual environment.

[0111] As a concrete example, if a user visits a virtual history museum and asks, "Who are the major artists of the Renaissance?", the server will use AI to instantly generate an answer such as, "Michelangelo and Leonardo da Vinci are representative artists of the Renaissance," and provide it in voice. This process utilizes Google's speech recognition API and a question-answering model from the transformers library.

[0112] The following are examples of input prompts for a generative AI model.

[0113] Question: Who were the major artists of the Renaissance?

[0114] Context: Information about Renaissance art and history.

[0115] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0116] Step 1:

[0117] The user enters a virtual environment and requests an audio guide on a specific topic. The terminal uses data storage from a server to retrieve relevant digital learning materials, which triggers the use of the audio generation system.

[0118] Step 2:

[0119] The server utilizes speech generation technology to create audio data from digital learning materials. The input consists of user-specified topic information and related context data. The output is audio data. A program within the server encodes text data into audio data.

[0120] Step 3:

[0121] The server transfers the generated audio data to the terminal, which then plays the audio using its audio output device. By listening to this, the user receives guide information within the virtual environment.

[0122] Step 4:

[0123] The user uses a voice input system to verbally request further information or ask specific questions to the terminal. This speech is transmitted to the terminal as voice input. Voice recognition software converts the input into text data and sends it to the server as query data.

[0124] Step 5:

[0125] The server uses an answer generation mechanism to generate immediate responses based on text data and leveraging an AI model. User inquiry data and digital learning materials serve as input. The output is a clear, textual answer to the user's question. The AI ​​model generates question-answers using contextually relevant prompts.

[0126] Step 6:

[0127] The server converts the response text back into audio data and sends it to the terminal. The terminal then plays the generated response aloud to the user using its audio output device. This allows the user to resolve any questions they may have on the spot.

[0128] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0129] This invention is a system for conducting lessons based on digital educational materials, and in particular, incorporates an emotion engine that recognizes the user's emotional state and adjusts the lesson content and responses accordingly. This provides a more interactive and personalized educational experience compared to conventional systems.

[0130] Specifically, the server stores digital educational materials as information storage and supplies them to the generation means as audio data necessary for the lesson to progress. The generation means generates appropriate audio for the lesson topic selected by the user and plays it back through the audio output means. When the audio is played back, an emotion engine analyzes the user's emotional state and adjusts the tone and speed of the audio based on the results.

[0131] The terminal conducts lessons through an audio output device and accepts user input as needed. During lessons, users communicate questions and feedback from students to the terminal, which is then converted into digital data by the input reception device and sent to the server.

[0132] The server generates an answer using a response generation mechanism based on the received question data. Furthermore, the emotion engine adjusts the tone and content of the answer according to the user's emotional state. For example, if the user appears confused, the server generates a response that includes a detailed explanation or a simple example. The generated answer is then converted back into audio data and played back through the terminal using an audio output mechanism.

[0133] This allows users to provide direct instructions and additional supplementary information, which is expected to improve the quality of lessons. By incorporating an emotion engine, the system can understand the user's emotions in real time and provide an optimal lesson experience. For example, if a student asks a difficult question and the user shows a confused expression, the emotion engine will recognize this state and the server will adjust the voice to provide a simpler and easier-to-understand explanation. This allows the user to continue the lesson efficiently.

[0134] The following describes the processing flow.

[0135] Step 1:

[0136] The server stores digital educational materials in an information storage device and makes them accessible. The server prepares to supply the materials to the generation device as needed.

[0137] Step 2:

[0138] The user operates their device to start a lesson and select a lesson topic. This information is sent from the device to the server as a request.

[0139] Step 3:

[0140] The server selects educational materials corresponding to the lesson topic and generates audio data using a generation method. The server then sends this data to the terminal.

[0141] Step 4:

[0142] The terminal plays the audio data received from the server using the audio output device, and the lesson begins. At this time, the emotion engine analyzes the user's facial expressions and tone of voice to determine their emotional state.

[0143] Step 5:

[0144] The emotion engine analyzes the user's emotional state in real time and adjusts the tone and speed of the audio during the lesson based on the information obtained. For example, if the user appears relaxed, the tone will be made calmer.

[0145] Step 6:

[0146] When a student has a question, the user receives the question using the input method on their device. This data is then converted into a digital format and sent to the server.

[0147] Step 7:

[0148] The server receives the question data and generates the optimal answer using a response generation mechanism. The server adjusts the answer based on feedback from the emotion engine; for example, if it determines that the user is confused, it will include a clearer explanation.

[0149] Step 8:

[0150] The terminal plays back the audio data of the answer received from the server using the response output device, providing the answer to the student. This allows the student to obtain a solution to their question.

[0151] Step 9:

[0152] The user monitors the lesson and supports student understanding by adjusting the audio in real time and providing additional explanations as needed. Finally, when the lesson ends, the system is shut down via the terminal.

[0153] (Example 2)

[0154] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0155] Traditional education systems often conducted lessons unilaterally without considering students' emotional states, making it difficult to provide an optimized learning experience for each individual learner. Furthermore, they were unable to immediately address students' confusion and anxiety when understanding complex material, resulting in decreased learning efficiency.

[0156] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0157] In this invention, the server includes means for storing information, processing means for receiving data from the information and generating speech, and means for outputting the speech generated by the processing means. This enables real-time analysis of the learner's emotional state and adjustment of speech tone and speed based on the results. This provides a personalized and interactive educational experience and improves the quality of learning.

[0158] "Means for storing information" refers to devices or systems that store digital educational materials and related data on media such as storage devices, and make them accessible as needed.

[0159] "A means of generating speech" refers to a device or program that creates synthesized speech from text data or other information.

[0160] "Means for outputting sound" refers to a device or interface for playing back the generated sound through speakers or headphones.

[0161] "An input means that receives voice input from participants and generates data" refers to a device or system that captures participants' speech using an input device such as a microphone, analyzes it, and converts it into digital data.

[0162] "Generating means for creating voice responses" refers to a program or system that generates appropriate response messages based on the participant's questions and feedback.

[0163] "Output means for reproducing a voice response" refers to a device or system that reproduces the generated voice response through an audio output device.

[0164] "An adjustment method that analyzes the emotional state of the participant and adjusts the tone and speed of the voice" refers to a device or system that judges the participant's emotions from their facial expressions and voice, and dynamically changes the characteristics of the generated voice based on that information.

[0165] A "correction mechanism for modifying response content based on emotions" refers to a device or program that appropriately modifies the content or expression of an existing response, taking into account the emotional state of the participant.

[0166] This invention is a lesson progression system based on digital educational materials that analyzes the emotional state of students in real time and appropriately adjusts the lesson content to provide an individualized educational experience. Specifically, the system is configured as follows.

[0167] The server stores digital educational materials using information storage means. A database is used for managing and providing educational materials. When a lesson begins, the server selects relevant materials according to the lesson topic chosen by the user and generates them as audio data using processing means. In this process, text-to-speech software or a speech synthesis engine is used to generate high-quality audio from the text.

[0168] The device plays back audio generated through its speaker via an audio output device, and also monitors the participant's facial expressions and behavior using its built-in camera and sensors. This allows the device to continuously collect and analyze data necessary for the emotion engine. If the participant shows signs of confusion or misunderstanding, the device sends instructions to the server to adjust the tone and pace of the audio. It also accepts voice input from the user, generates question data, and sends it to the server.

[0169] When a student asks a question, the server uses a response generation mechanism to return an appropriate answer to the student. Utilizing a generative AI model, it creates detailed and easy-to-understand answers based on the question, with an emotion engine flexibly adjusting the tone and content as needed. The generated answer is then converted back into audio data and provided to the student via their device.

[0170] For example, if a student asks a question about a scientific topic and shows signs of confusion, the server may generate a modified response such as, "This concept is often perceived as difficult, so let me explain it in more detail." An example of a prompt would be, "If a user shows signs of confusion during class, simplify the explanation and soften the tone."

[0171] This system, with its dynamic content adjustments tailored to each student, is expected to improve learning efficiency and maximize comprehension.

[0172] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0173] Step 1:

[0174] The server retrieves digital educational materials related to the user's selected lesson topic from its information storage device. The input is the selected lesson topic information, and the output is the corresponding lesson material data. The lesson material data is retrieved as a text file.

[0175] Step 2:

[0176] The server converts the acquired educational material data into audio data using processing tools. The input is text-based educational material data, and the output is a synthesized audio file. Specifically, text-to-speech software is used to convert the text into speech.

[0177] Step 3:

[0178] The terminal receives audio data transmitted from the server and plays it back to the learner through an audio output device. The input is an audio file, and the output is audio playback to the learner. Specifically, the terminal plays the generated audio through its speaker.

[0179] Step 4:

[0180] The device uses cameras and sensors to collect real-time facial expression data from participants. Visual data of the participants is the input, and facial expression data is obtained as the output. Specifically, facial recognition software is used to analyze the participants' emotional state.

[0181] Step 5:

[0182] The emotion engine determines the learner's emotional state based on acquired facial expression data and suggests adjustments to voice tone and speed. The input is facial expression analysis data, and the output is adjustment instructions. Based on this, the audio being played is dynamically adjusted.

[0183] Step 6:

[0184] The user enters a question into the terminal, and the input receiving device sends it to the server as voice or text data. The input is the student's question data, and the output is sent to the server in digital format. Specifically, speech recognition software is used to convert the voice into text.

[0185] Step 7:

[0186] The server generates appropriate answers using a generative AI model based on the received question data. The input is the question data, and the output is a speech-convertible answer text. Specifically, the AI ​​model presents information from a relevant knowledge base.

[0187] Step 8:

[0188] The generated response is converted into audio data, sent to the terminal, and played back by the audio output device. The input is response text data, and the output is an audio response. Specifically, a speech synthesis engine is used to generate and play back a smooth response.

[0189] This series of steps enables interactive and personalized education that responds to the emotional state of the participants.

[0190] (Application Example 2)

[0191] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0192] In modern education, a challenge is providing an educational environment that is tailored to each learner's individual level of understanding and emotional state. Traditional educational systems have struggled to customize lesson progression and responses to accommodate learners' emotions, making it difficult to provide optimal education that meets individual needs. Especially in online and virtual environments, there is a need for mechanisms that maintain learners' interest and concentration and promote effective learning.

[0193] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0194] In this invention, the server includes an information storage means for providing digital learning materials, an emotion analysis means for analyzing the learner's emotional state, and an educational content adjustment means for dynamically adjusting the educational content based on the analysis results. This makes it possible to provide interactive and personalized educational content that responds to the learner's emotional state.

[0195] "Digital learning materials" refer to educational content and teaching materials that are available on computer systems and digital devices and are stored electronically.

[0196] "Information storage means" refers to physical or electronic devices or systems for storing digital learning materials and related information.

[0197] "Generation means" refers to a mechanism or process for forming content such as audio based on data received from information storage means.

[0198] "Audio output means" refers to a device or system used to allow a user to hear the generated audio.

[0199] "Input receiving means" refers to a device or method that receives voice input from learners and processes it as necessary information.

[0200] "Question data" refers to formalized information containing questions and inquiries that learners raise during class.

[0201] "Response generation means" refers to a system or method for constructing appropriate answers or responses based on question data.

[0202] "Response output means" refers to a device or function for transmitting the generated response to the user.

[0203] "Emotional analysis methods" refer to techniques and technologies used to analyze a user's emotional state and adjust the system's operation based on that information.

[0204] "Educational content adjustment means" refers to a mechanism that appropriately adjusts the educational content provided based on the results of an analysis of emotional states.

[0205] To implement this invention, the program must be installed on both the server and the user's terminal. The server stores digital learning materials in an information storage means and generates appropriate audio based on this information using a generation means. The generated audio is output to the user's terminal via an audio output means.

[0206] The user's terminal receives voice input from the learner using an input receiving means, generates question data, and sends it to the server. The server uses a response generation means to create an answer based on the question data. At that time, an emotion analysis means analyzes the user's emotional state from facial expressions and other factors, and adjusts the tone and content of the response based on the results.

[0207] The voice response is played back by the response output device, and the educational content adjustment device dynamically modifies the learning materials according to the analyzed emotion. This makes adjustments that allow learners to understand more effectively.

[0208] For example, if a user shows a confused expression regarding an assignment they are working on, the emotion analysis tool will detect this. Based on the analysis results, the server adjusts the educational content to add simple and visually easy-to-understand examples. For instance, a prompt such as, "If a student is bored, suggest how to make learning more engaging," is used to have the AI ​​model generate content. In this way, by utilizing prompts generated by a generative AI model, a customized educational experience can be provided to the user.

[0209] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0210] Step 1:

[0211] The server retrieves digital learning materials from the information storage device. If the information is recorded as text, images, or audio data, it identifies the necessary learning materials and converts them to the appropriate format. The input is data from the information storage device, and the output is learning material data for the generation device. The server organizes the digital materials and prepares them for generation.

[0212] Step 2:

[0213] The server uses a generation mechanism to generate audio from the received educational material data. During this process, speech synthesis technology is used to create content with an easy-to-understand volume and tone, depending on the content of the material. The input is the educational material data, and the output is audio data. The server executes the speech synthesis process and prepares the audio for transmission to the terminal.

[0214] Step 3:

[0215] The terminal plays the audio data received via its audio output device. This operation is performed through the terminal's speaker, enabling the user to begin learning. The input is audio data from the server, and the output is the audio the user hears. The terminal plays the audio and provides information to the user.

[0216] Step 4:

[0217] The user communicates opinions and questions to the terminal via voice input through an input receiving device. This data is digitized by the terminal and generated as question data. The input is the user's voice, and the output is question data sent to the server. The user clearly expresses their question, and the terminal creates the necessary data.

[0218] Step 5:

[0219] The server receives question data and analyzes it using a response generation system. It generates an answer that matches the content of the question and also considers the user's emotional state using an emotion analysis system. Based on this information, the server adjusts the content of the answer. The input is question data and emotion information, and the output is the adjusted response data. The server dynamically generates answers to improve relevance.

[0220] Step 6:

[0221] The terminal uses a response output mechanism to play back the response data provided by the server as audio. At this stage, the learner receives feedback from the server. The input is the response data from the server, and the output is the audio the learner hears. The terminal converts the answer into audio, facilitating smooth communication with the user.

[0222] Step 7:

[0223] The server activates an educational content adjustment mechanism based on information obtained from sentiment analysis, and reconfigures learning materials as needed. It utilizes a generative AI model to generate new educational content using prompts, changing the next learning objective. Input is the user's sentiment data and prompts, and output is the adjusted learning content. The server continuously optimizes the educational content to support the user's learning.

[0224] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0225] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0226] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0227] [Second Embodiment]

[0228] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0229] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0230] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0231] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0232] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0233] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0234] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0235] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0236] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0237] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0238] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0239] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0240] This invention is a system for managing lesson progress and responding to student questions based on digitized educational materials. The following describes a specific implementation of this system.

[0241] First, at the heart of the system resides digital educational materials, and the information storage means that stores these materials acts as a server. The server accesses the digital educational materials and, based on this, supplies the audio necessary for the lesson to the audio generation means. The server provides the necessary educational content according to the data requested by the audio generation means and distributes the generated audio data to the terminals.

[0242] Next, the terminal plays the audio data received from the server and conducts the classroom lesson through the audio output device. The terminal selects appropriate audio according to the context of the lesson and supports the flow of the lesson. The terminal is also equipped with an input device to receive questions from students via voice. The questions entered via voice are converted into digital data and sent to the server.

[0243] The server generates appropriate answers from digital educational materials based on the received question data. To do this, the server uses response generation means and AI algorithms to quickly construct appropriate answers. The answers are then sent back to the terminal as audio data.

[0244] The terminal provides responses to student questions by playing back audio data received from the server using its response output mechanism. This allows for immediate feedback on student inquiries without interrupting the flow of the lesson.

[0245] The user (teacher) can monitor the progress of the lesson, intervene in the device as needed, and pause or add explanations to the audio. This allows the teacher to confirm students' understanding and take on roles such as supplementing key points.

[0246] As a concrete example, when lecturing on "Edo period culture" in a history class, the server generates audio for the relevant topic from a digitized history textbook, and the terminals play this audio in the classroom. If a student asks, "I want to know more about Edo period clothing," the terminal transmits this as voice input to the server, which uses a response generation mechanism to provide a detailed explanation and returns it to the student as voice output.

[0247] Thus, this system is designed to efficiently support lessons while reducing the burden on teachers.

[0248] The following describes the processing flow.

[0249] Step 1:

[0250] The server stores digital educational materials in an information storage device and makes them accessible as needed. The server organizes teaching material data related to topics used in lessons and prepares it so that the generation device can efficiently generate audio.

[0251] Step 2:

[0252] At the start of a class, the user operates their device and selects a class topic. This action causes the device to send a request to the server based on the selected topic and begin retrieving the necessary educational materials.

[0253] Step 3:

[0254] The server, in response to the user's request, passes educational materials on the selected topic to the audio generation device. The server then converts the specified content into audio data and prepares it for transmission to the terminal.

[0255] Step 4:

[0256] The terminal plays back audio data received from the server using an audio output device. This allows the lesson to begin with an audio guide such as, "Today we will learn about XX."

[0257] Step 5:

[0258] When a student asks a question, the terminal's input receiving mechanism receives it as voice input and converts it into digital data. The terminal then sends the converted question data to the server and requests the generation of a response.

[0259] Step 6:

[0260] The server receives the question data and uses a response generation mechanism to create an answer based on digital educational materials. The server generates the optimal answer to the question and sends it to the terminal as audio data.

[0261] Step 7:

[0262] The device plays back the received audio data of the response via a response output device, providing the student with the answer to their question. This allows students to resolve their questions without disrupting the flow of the lesson.

[0263] Step 8:

[0264] The user monitors the lesson progress and controls audio playback via their device as needed. The user can provide supplementary explanations and adjust the lesson pace based on student reactions.

[0265] Step 9:

[0266] When the lesson ends, the user exits the lesson by operating their terminal. The server creates a lesson log and stores data to help improve future lessons.

[0267] (Example 1)

[0268] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0269] In today's educational environment, teachers spend a great deal of time and effort managing the progress of lessons and responding to individual student questions. As a result, there is a risk of insufficient time for home education and a decline in the quality of lesson content. In particular, the burden on teachers increases when individualized support tailored to each student's level of understanding is required. To address these challenges, a system is needed that allows for efficient lesson progression and immediate responses to individual student questions.

[0270] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0271] In this invention, the server includes a storage device means for providing digital teaching materials, a speech synthesis device means for acquiring digital teaching materials from the storage device and generating audio, and an input device means for receiving voice input from learners and generating inquiry data. This makes it possible to respond to students' questions in real time while smoothly conducting lessons.

[0272] A "storage device" is a device used to store and keep digital learning materials accessible.

[0273] A "speech synthesis device" is a device that generates audio data from information acquired from digital educational materials.

[0274] A "playback device" is a device that outputs generated audio data and allows learners to listen to it.

[0275] An "input device" is a device that receives voice input from learners and processes it as inquiry data.

[0276] A "response generation device" is a device that generates an appropriate voice response based on inquiry data.

[0277] An "output device" is a device that plays back the generated voice response and provides it to the learner.

[0278] A "generative artificial intelligence model" is a computational model that utilizes information technology to generate optimized responses based on query data.

[0279] This invention relates to an educational support system that utilizes digital teaching materials, and is configured in which a server, a terminal, and a user work together in a coordinated manner.

[0280] The server is the core of the education support system. It stores digital teaching materials in cloud storage and functions as a storage device. When a request for acquiring teaching materials occurs, the server uses a voice synthesis device to utilize voice technologies such as the Google Text-to-Speech API and converts the teaching materials into voice data. The converted voice data is streamed to the terminal.

[0281] The terminal is a device that plays the voice data received from the server. The terminal has a playback device and supports the progress of the class by outputting voice in the classroom. Also, an input device is provided on the terminal, which is responsible for receiving voice input from the learner and sending it to the server. The input questions are appropriately processed and converted into digital data.

[0282] The user (teacher), as the supervisor of the system, can pause the voice output temporarily or provide additional explanations via the terminal during the class as needed. This enables flexible class progress based on the learners' understanding. Also, the user checks the questions from the learners and provides supplementary information for areas that require further exploration.

[0283] As a specific example, consider the scenario of teaching "Culture in Medieval Europe" in a history class. The server generates the voice of the corresponding topic from the digitized history teaching materials, and the terminal plays that voice in the classroom. For example, when a learner asks "Please tell me about the clothing in medieval Europe", the terminal transfers the question to the server using voice recognition technology. The server uses a response generation device and drives the generation AI model to generate an appropriate answer. This answer is converted back into voice data and provided to the learner via the terminal.

[0284] As a specific example of the prompt text, by inputting "Please explain the clothing in medieval Europe in detail", the generation AI model can extract relevant information and provide a detailed explanation. In this way, the present invention realizes the improvement of efficiency in the educational field and two-way learning support.

[0285] The flow of the specific process in Example 1 will be described using FIG. 11.

[0286] Step 1:

[0287] The server acquires the digital teaching materials from the cloud storage and stores them in the storage device. This input includes the file format of the digital teaching materials. Upon receiving the request for teaching material acquisition, the server searches for the corresponding content and prepares the data. The output is a dataset in which the teaching material content awaits in an accessible state.

[0288] Step 2:

[0289] The server uses a speech synthesis device to perform processing in order to convert the teaching material data into speech data. As a specific operation, it utilizes speech technologies such as the Google Text-to-Speech API to convert the text data into a speech file. In this process, the text data is input into the speech synthesis device and output as speech data.

[0290] Step 3:

[0291] The server streams and distributes the generated speech data to the terminal. As a result, the speech for use in the class is delivered in real time. The input is the converted speech data, and the output is the speech streaming distribution to the terminal via the network.

[0292] Step 4:

[0293] The terminal outputs the speech data received from the server in the classroom by means of a playback device. For the start of playback, it is necessary for the speech data to be input into the terminal. As a specific operation, the speech is played back throughout the classroom through the speaker. The output is the speech audible to the class participants.

[0294] Step 5:

[0295] The terminal receives voice input from the learner and converts it into query data using an input device. The input includes the learner's question, which is collected via a microphone and processed into digital data. The resulting output is query data that is sent to the server.

[0296] Step 6:

[0297] The server processes the query data using a response generator and generates an appropriate response using a generation AI model. This process involves designing a prompt and inputting it into the model to obtain a response in appropriate text format. The output is the response information that should be converted into audio data.

[0298] Step 7:

[0299] The server converts the generated text responses into speech and then sends them to the terminal. A text-to-speech synthesis device is used for this conversion. The input is the response generated by the AI ​​model, and the output is the audio data streamed to the terminal.

[0300] Step 8:

[0301] The terminal finally outputs the audio data received from the server back into the classroom via a playback device. This completes the feedback process to the learners. The input is audio data, and the output is an audible response.

[0302] (Application Example 1)

[0303] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0304] In the current education system, it is difficult to provide an interactive learning experience in the classroom, especially to obtain quick and accurate feedback in real-time question and answer sessions. Also, when exploring information in a virtual environment, the voice guides and explanations received by users are limited, leaving much room for expanding the depth of learning. As a result, issues such as providing an efficient learning environment and the quality of education being restricted are raised.

[0305] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following respective means.

[0306] In this invention, the server includes a data storage means, a voice generation means, and a voice output means. Thereby, users can search for information in real-time in a virtual environment and receive an immediate response by AI.

[0307] The "data storage means" is a device that stores digital learning materials and can retrieve them as needed.

[0308] The "voice generation means" is a device that creates voice data based on the stored digital learning materials.

[0309] The "voice output means" is a device that physically reproduces the generated voice data and provides it to the user.

[0310] The "voice reception means" is a device that recognizes voice input from the user and converts it into digital data.

[0311] The "answer generation means" is a device that constructs appropriate information in real-time based on inquiries from the user and returns an answer.

[0312] The "answer output means" is a device that reproduces the generated answer and conveys it to the user.

[0313] An "exploration tool" is a device that provides guidance to users when they search for and learn information in a virtual space.

[0314] An "interactive response system" is a device that uses AI to enable immediate responses to user questions.

[0315] To implement this invention, a system is configured to support learning and exploration in a virtual environment. The server stores digital learning materials and uses a speech generation means to create audio data related to the information requested by the user. The speech output means transmits this audio data to a terminal and plays it back in a format audible to the user. The terminal is equipped with a speech receiving means that receives the user's voice input, converts it into digital data, and transmits it to the server.

[0316] The server utilizes response generation mechanisms and constructs instant responses using AI algorithms. The generated responses are then converted back into audio data and delivered to the user through response output mechanisms. This allows users to gain an interactive learning experience in a virtual environment.

[0317] As a concrete example, if a user visits a virtual history museum and asks, "Who are the major artists of the Renaissance?", the server will use AI to instantly generate an answer such as, "Michelangelo and Leonardo da Vinci are representative artists of the Renaissance," and provide it in voice. This process utilizes Google's speech recognition API and a question-answering model from the transformers library.

[0318] The following are examples of input prompts for a generative AI model.

[0319] Question: Who were the major artists of the Renaissance?

[0320] Context: Information about Renaissance art and history.

[0321] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0322] Step 1:

[0323] The user enters a virtual environment and requests an audio guide on a specific topic. The terminal uses data storage from a server to retrieve relevant digital learning materials, which triggers the use of the audio generation system.

[0324] Step 2:

[0325] The server utilizes speech generation technology to create audio data from digital learning materials. The input consists of user-specified topic information and related context data. The output is audio data. A program within the server encodes text data into audio data.

[0326] Step 3:

[0327] The server transfers the generated audio data to the terminal, which then plays the audio using its audio output device. By listening to this, the user receives guide information within the virtual environment.

[0328] Step 4:

[0329] The user uses a voice input system to verbally request further information or ask specific questions to the terminal. This speech is transmitted to the terminal as voice input. Voice recognition software converts the input into text data and sends it to the server as query data.

[0330] Step 5:

[0331] The server uses an answer generation mechanism to generate immediate responses based on text data and leveraging an AI model. User inquiry data and digital learning materials serve as input. The output is a clear, textual answer to the user's question. The AI ​​model generates question-answers using contextually relevant prompts.

[0332] Step 6:

[0333] The server converts the response text back into audio data and sends it to the terminal. The terminal then plays the generated response aloud to the user using its audio output device. This allows the user to resolve any questions they may have on the spot.

[0334] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0335] This invention is a system for conducting lessons based on digital educational materials, and in particular, incorporates an emotion engine that recognizes the user's emotional state and adjusts the lesson content and responses accordingly. This provides a more interactive and personalized educational experience compared to conventional systems.

[0336] Specifically, the server stores digital educational materials as information storage and supplies them to the generation means as audio data necessary for the lesson to progress. The generation means generates appropriate audio for the lesson topic selected by the user and plays it back through the audio output means. When the audio is played back, an emotion engine analyzes the user's emotional state and adjusts the tone and speed of the audio based on the results.

[0337] The terminal conducts lessons through an audio output device and accepts user input as needed. During lessons, users communicate questions and feedback from students to the terminal, which is then converted into digital data by the input reception device and sent to the server.

[0338] The server generates an answer using a response generation mechanism based on the received question data. Furthermore, the emotion engine adjusts the tone and content of the answer according to the user's emotional state. For example, if the user appears confused, the server generates a response that includes a detailed explanation or a simple example. The generated answer is then converted back into audio data and played back through the terminal using an audio output mechanism.

[0339] This allows users to provide direct instructions and additional supplementary information, which is expected to improve the quality of lessons. By incorporating an emotion engine, the system can understand the user's emotions in real time and provide an optimal lesson experience. For example, if a student asks a difficult question and the user shows a confused expression, the emotion engine will recognize this state and the server will adjust the voice to provide a simpler and easier-to-understand explanation. This allows the user to continue the lesson efficiently.

[0340] The following describes the processing flow.

[0341] Step 1:

[0342] The server stores digital educational materials in an information storage device and makes them accessible. The server prepares to supply the materials to the generation device as needed.

[0343] Step 2:

[0344] The user operates their device to start a lesson and select a lesson topic. This information is sent from the device to the server as a request.

[0345] Step 3:

[0346] The server selects educational materials corresponding to the lesson topic and generates audio data using a generation method. The server then sends this data to the terminal.

[0347] Step 4:

[0348] The terminal plays the audio data received from the server using the audio output device, and the lesson begins. At this time, the emotion engine analyzes the user's facial expressions and tone of voice to determine their emotional state.

[0349] Step 5:

[0350] The emotion engine analyzes the user's emotional state in real time and adjusts the tone and speed of the audio during the lesson based on the information obtained. For example, if the user appears relaxed, the tone will be made calmer.

[0351] Step 6:

[0352] When a student has a question, the user receives the question using the input method on their device. This data is then converted into a digital format and sent to the server.

[0353] Step 7:

[0354] The server receives the question data and generates the optimal answer using a response generation mechanism. The server adjusts the answer based on feedback from the emotion engine; for example, if it determines that the user is confused, it will include a clearer explanation.

[0355] Step 8:

[0356] The terminal plays back the audio data of the answer received from the server using the response output device, providing the answer to the student. This allows the student to obtain a solution to their question.

[0357] Step 9:

[0358] The user monitors the lesson and supports student understanding by adjusting the audio in real time and providing additional explanations as needed. Finally, when the lesson ends, the system is shut down via the terminal.

[0359] (Example 2)

[0360] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0361] Traditional education systems often conducted lessons unilaterally without considering students' emotional states, making it difficult to provide an optimized learning experience for each individual learner. Furthermore, they were unable to immediately address students' confusion and anxiety when understanding complex material, resulting in decreased learning efficiency.

[0362] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0363] In this invention, the server includes means for storing information, processing means for receiving data from the information and generating speech, and means for outputting the speech generated by the processing means. This enables real-time analysis of the learner's emotional state and adjustment of speech tone and speed based on the results. This provides a personalized and interactive educational experience and improves the quality of learning.

[0364] "Means for storing information" refers to devices or systems that store digital educational materials and related data on media such as storage devices, and make them accessible as needed.

[0365] "A means of generating speech" refers to a device or program that creates synthesized speech from text data or other information.

[0366] "Means for outputting sound" refers to a device or interface for playing back the generated sound through speakers or headphones.

[0367] "An input means that receives voice input from participants and generates data" refers to a device or system that captures participants' speech using an input device such as a microphone, analyzes it, and converts it into digital data.

[0368] "Generating means for creating voice responses" refers to a program or system that generates appropriate response messages based on the participant's questions and feedback.

[0369] "Output means for reproducing a voice response" refers to a device or system that reproduces the generated voice response through an audio output device.

[0370] "An adjustment method that analyzes the emotional state of the participant and adjusts the tone and speed of the voice" refers to a device or system that judges the participant's emotions from their facial expressions and voice, and dynamically changes the characteristics of the generated voice based on that information.

[0371] A "correction mechanism for modifying response content based on emotions" refers to a device or program that appropriately modifies the content or expression of an existing response, taking into account the emotional state of the participant.

[0372] This invention is a lesson progression system based on digital educational materials that analyzes the emotional state of students in real time and appropriately adjusts the lesson content to provide an individualized educational experience. Specifically, the system is configured as follows.

[0373] The server stores digital educational materials using information storage means. A database is used for managing and providing educational materials. When a lesson begins, the server selects relevant materials according to the lesson topic chosen by the user and generates them as audio data using processing means. In this process, text-to-speech software or a speech synthesis engine is used to generate high-quality audio from the text.

[0374] The device plays back audio generated through its speaker via an audio output device, and also monitors the participant's facial expressions and behavior using its built-in camera and sensors. This allows the device to continuously collect and analyze data necessary for the emotion engine. If the participant shows signs of confusion or misunderstanding, the device sends instructions to the server to adjust the tone and pace of the audio. It also accepts voice input from the user, generates question data, and sends it to the server.

[0375] When a student asks a question, the server uses a response generation mechanism to return an appropriate answer to the student. Utilizing a generative AI model, it creates detailed and easy-to-understand answers based on the question, with an emotion engine flexibly adjusting the tone and content as needed. The generated answer is then converted back into audio data and provided to the student via their device.

[0376] For example, if a student asks a question about a scientific topic and shows signs of confusion, the server may generate a modified response such as, "This concept is often perceived as difficult, so let me explain it in more detail." An example of a prompt would be, "If a user shows signs of confusion during class, simplify the explanation and soften the tone."

[0377] This system, with its dynamic content adjustments tailored to each student, is expected to improve learning efficiency and maximize comprehension.

[0378] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0379] Step 1:

[0380] The server retrieves digital educational materials related to the user's selected lesson topic from its information storage device. The input is the selected lesson topic information, and the output is the corresponding lesson material data. The lesson material data is retrieved as a text file.

[0381] Step 2:

[0382] The server converts the acquired educational material data into audio data using processing tools. The input is text-based educational material data, and the output is a synthesized audio file. Specifically, text-to-speech software is used to convert the text into speech.

[0383] Step 3:

[0384] The terminal receives audio data transmitted from the server and plays it back to the learner through an audio output device. The input is an audio file, and the output is audio playback to the learner. Specifically, the terminal plays the generated audio through its speaker.

[0385] Step 4:

[0386] The device uses cameras and sensors to collect real-time facial expression data from participants. Visual data of the participants is the input, and facial expression data is obtained as the output. Specifically, facial recognition software is used to analyze the participants' emotional state.

[0387] Step 5:

[0388] The emotion engine determines the learner's emotional state based on acquired facial expression data and suggests adjustments to voice tone and speed. The input is facial expression analysis data, and the output is adjustment instructions. Based on this, the audio being played is dynamically adjusted.

[0389] Step 6:

[0390] The user enters a question into the terminal, and the input receiving device sends it to the server as voice or text data. The input is the student's question data, and the output is sent to the server in digital format. Specifically, speech recognition software is used to convert the voice into text.

[0391] Step 7:

[0392] The server generates appropriate answers using a generative AI model based on the received question data. The input is the question data, and the output is a speech-convertible answer text. Specifically, the AI ​​model presents information from a relevant knowledge base.

[0393] Step 8:

[0394] The generated response is converted into audio data, sent to the terminal, and played back by the audio output device. The input is response text data, and the output is an audio response. Specifically, a speech synthesis engine is used to generate and play back a smooth response.

[0395] This series of steps enables interactive and personalized education that responds to the emotional state of the participants.

[0396] (Application Example 2)

[0397] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the smart glasses 214 as the "terminal".

[0398] In modern education, a challenge is providing an educational environment that is tailored to each learner's individual level of understanding and emotional state. Traditional educational systems have struggled to customize lesson progression and responses to accommodate learners' emotions, making it difficult to provide optimal education that meets individual needs. Especially in online and virtual environments, there is a need for mechanisms that maintain learners' interest and concentration and promote effective learning.

[0399] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0400] In this invention, the server includes an information storage means for providing digital learning materials, an emotion analysis means for analyzing the learner's emotional state, and an educational content adjustment means for dynamically adjusting the educational content based on the analysis results. This makes it possible to provide interactive and personalized educational content that responds to the learner's emotional state.

[0401] "Digital learning materials" refer to educational content and teaching materials that are available on computer systems and digital devices and are stored electronically.

[0402] "Information storage means" refers to physical or electronic devices or systems for storing digital learning materials and related information.

[0403] "Generation means" refers to a mechanism or process for forming content such as audio based on data received from information storage means.

[0404] "Audio output means" refers to a device or system used to allow a user to hear the generated audio.

[0405] "Input receiving means" refers to a device or method that receives voice input from learners and processes it as necessary information.

[0406] "Question data" refers to formalized information containing questions and inquiries that learners raise during class.

[0407] "Response generation means" refers to a system or method for constructing appropriate answers or responses based on question data.

[0408] "Response output means" refers to a device or function for transmitting the generated response to the user.

[0409] "Emotional analysis methods" refer to techniques and technologies used to analyze a user's emotional state and adjust the system's operation based on that information.

[0410] "Educational content adjustment means" refers to a mechanism that appropriately adjusts the educational content provided based on the results of an analysis of emotional states.

[0411] To implement this invention, the program must be installed on both the server and the user's terminal. The server stores digital learning materials in an information storage means and generates appropriate audio based on this information using a generation means. The generated audio is output to the user's terminal via an audio output means.

[0412] The user's terminal receives voice input from the learner using an input receiving means, generates question data, and sends it to the server. The server uses a response generation means to create an answer based on the question data. At that time, an emotion analysis means analyzes the user's emotional state from facial expressions and other factors, and adjusts the tone and content of the response based on the results.

[0413] The voice response is played back by the response output device, and the educational content adjustment device dynamically modifies the learning materials according to the analyzed emotion. This makes adjustments that allow learners to understand more effectively.

[0414] For example, if a user shows a confused expression regarding an assignment they are working on, the emotion analysis tool will detect this. Based on the analysis results, the server adjusts the educational content to add simple and visually easy-to-understand examples. For instance, a prompt such as, "If a student is bored, suggest how to make learning more engaging," is used to have the AI ​​model generate content. In this way, by utilizing prompts generated by a generative AI model, a customized educational experience can be provided to the user.

[0415] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0416] Step 1:

[0417] The server retrieves digital learning materials from the information storage device. If the information is recorded as text, images, or audio data, it identifies the necessary learning materials and converts them to the appropriate format. The input is data from the information storage device, and the output is learning material data for the generation device. The server organizes the digital materials and prepares them for generation.

[0418] Step 2:

[0419] The server uses a generation mechanism to generate audio from the received educational material data. During this process, speech synthesis technology is used to create content with an easy-to-understand volume and tone, depending on the content of the material. The input is the educational material data, and the output is audio data. The server executes the speech synthesis process and prepares the audio for transmission to the terminal.

[0420] Step 3:

[0421] The terminal plays the audio data received via its audio output device. This operation is performed through the terminal's speaker, enabling the user to begin learning. The input is audio data from the server, and the output is the audio the user hears. The terminal plays the audio and provides information to the user.

[0422] Step 4:

[0423] The user communicates opinions and questions to the terminal via voice input through an input receiving device. This data is digitized by the terminal and generated as question data. The input is the user's voice, and the output is question data sent to the server. The user clearly expresses their question, and the terminal creates the necessary data.

[0424] Step 5:

[0425] The server receives question data and analyzes it using a response generation system. It generates an answer that matches the content of the question and also considers the user's emotional state using an emotion analysis system. Based on this information, the server adjusts the content of the answer. The input is question data and emotion information, and the output is the adjusted response data. The server dynamically generates answers to improve relevance.

[0426] Step 6:

[0427] The terminal uses a response output mechanism to play back the response data provided by the server as audio. At this stage, the learner receives feedback from the server. The input is the response data from the server, and the output is the audio the learner hears. The terminal converts the answer into audio, facilitating smooth communication with the user.

[0428] Step 7:

[0429] The server activates an educational content adjustment mechanism based on information obtained from sentiment analysis, and reconfigures learning materials as needed. It utilizes a generative AI model to generate new educational content using prompts, changing the next learning objective. Input is the user's sentiment data and prompts, and output is the adjusted learning content. The server continuously optimizes the educational content to support the user's learning.

[0430] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0431] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0432] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0433] [Third Embodiment]

[0434] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0435] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0436] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0437] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0438] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0439] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0440] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0441] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0442] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0443] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0444] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0445] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0446] This invention is a system for managing lesson progress and responding to student questions based on digitized educational materials. The following describes a specific implementation of this system.

[0447] First, at the heart of the system resides digital educational materials, and the information storage means that stores these materials acts as a server. The server accesses the digital educational materials and, based on this, supplies the audio necessary for the lesson to the audio generation means. The server provides the necessary educational content according to the data requested by the audio generation means and distributes the generated audio data to the terminals.

[0448] Next, the terminal plays the audio data received from the server and conducts the classroom lesson through the audio output device. The terminal selects appropriate audio according to the context of the lesson and supports the flow of the lesson. The terminal is also equipped with an input device to receive questions from students via voice. The questions entered via voice are converted into digital data and sent to the server.

[0449] The server generates appropriate answers from digital educational materials based on the received question data. To do this, the server uses response generation means and AI algorithms to quickly construct appropriate answers. The answers are then sent back to the terminal as audio data.

[0450] The terminal provides responses to student questions by playing back audio data received from the server using its response output mechanism. This allows for immediate feedback on student inquiries without interrupting the flow of the lesson.

[0451] The user (teacher) can monitor the progress of the lesson, intervene in the device as needed, and pause or add explanations to the audio. This allows the teacher to confirm students' understanding and take on roles such as supplementing key points.

[0452] As a concrete example, when lecturing on "Edo period culture" in a history class, the server generates audio for the relevant topic from a digitized history textbook, and the terminals play this audio in the classroom. If a student asks, "I want to know more about Edo period clothing," the terminal transmits this as voice input to the server, which uses a response generation mechanism to provide a detailed explanation and returns it to the student as voice output.

[0453] Thus, this system is designed to efficiently support lessons while reducing the burden on teachers.

[0454] The following describes the processing flow.

[0455] Step 1:

[0456] The server stores digital educational materials in an information storage device and makes them accessible as needed. The server organizes teaching material data related to topics used in lessons and prepares it so that the generation device can efficiently generate audio.

[0457] Step 2:

[0458] At the start of a class, the user operates their device and selects a class topic. This action causes the device to send a request to the server based on the selected topic and begin retrieving the necessary educational materials.

[0459] Step 3:

[0460] The server, in response to the user's request, passes educational materials on the selected topic to the audio generation device. The server then converts the specified content into audio data and prepares it for transmission to the terminal.

[0461] Step 4:

[0462] The terminal plays back audio data received from the server using an audio output device. This allows the lesson to begin with an audio guide such as, "Today we will learn about XX."

[0463] Step 5:

[0464] When a student asks a question, the terminal's input receiving mechanism receives it as voice input and converts it into digital data. The terminal then sends the converted question data to the server and requests the generation of a response.

[0465] Step 6:

[0466] The server receives the question data and uses a response generation mechanism to create an answer based on digital educational materials. The server generates the optimal answer to the question and sends it to the terminal as audio data.

[0467] Step 7:

[0468] The device plays back the received audio data of the response via a response output device, providing the student with the answer to their question. This allows students to resolve their questions without disrupting the flow of the lesson.

[0469] Step 8:

[0470] The user monitors the lesson progress and controls audio playback via their device as needed. The user can provide supplementary explanations and adjust the lesson pace based on student reactions.

[0471] Step 9:

[0472] When the lesson ends, the user exits the lesson by operating their terminal. The server creates a lesson log and stores data to help improve future lessons.

[0473] (Example 1)

[0474] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0475] In today's educational environment, teachers spend a great deal of time and effort managing the progress of lessons and responding to individual student questions. As a result, there is a risk of insufficient time for home education and a decline in the quality of lesson content. In particular, the burden on teachers increases when individualized support tailored to each student's level of understanding is required. To address these challenges, a system is needed that allows for efficient lesson progression and immediate responses to individual student questions.

[0476] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0477] In this invention, the server includes a storage device means for providing digital teaching materials, a speech synthesis device means for acquiring digital teaching materials from the storage device and generating audio, and an input device means for receiving voice input from learners and generating inquiry data. This makes it possible to respond to students' questions in real time while smoothly conducting lessons.

[0478] A "storage device" is a device used to store and keep digital learning materials accessible.

[0479] A "speech synthesis device" is a device that generates audio data from information acquired from digital educational materials.

[0480] A "playback device" is a device that outputs generated audio data and allows learners to listen to it.

[0481] An "input device" is a device that receives voice input from learners and processes it as inquiry data.

[0482] A "response generation device" is a device that generates an appropriate voice response based on inquiry data.

[0483] An "output device" is a device that plays back the generated voice response and provides it to the learner.

[0484] A "generative artificial intelligence model" is a computational model that utilizes information technology to generate optimized responses based on query data.

[0485] This invention relates to an educational support system that utilizes digital teaching materials, and is configured in which a server, a terminal, and a user work together in a coordinated manner.

[0486] The server is the core of the educational support system. It stores digital learning materials in cloud storage and functions as a memory device. When a request for learning materials is received, the server uses a speech synthesis device and utilizes speech technologies such as the Google Text-to-Speech API to convert the materials into audio data. The converted audio data is then streamed to the terminal.

[0487] The terminal is a device that plays audio data received from the server. Equipped with a playback device, the terminal supports the progress of lessons by outputting audio within the classroom. The terminal also has an input device that receives audio input from learners and transmits it to the server. The input questions are processed appropriately and converted into digital data.

[0488] The user (teacher), acting as the system's supervisor, pauses audio output and provides additional explanations via their terminal as needed during lessons. This allows for flexible lesson progression based on the learners' understanding. The user also reviews learners' questions and provides supplementary information where further explanation is required.

[0489] As a concrete example, consider a scenario in a history class where a lecture is given on "Medieval European Culture." The server generates audio for the relevant topic from digitized history materials, and the terminals play this audio in the classroom. For example, if a student asks, "Please tell me about medieval European clothing," the terminal uses speech recognition technology to transmit the question to the server. The server uses a response generation device and a generative AI model to generate an appropriate answer. This answer is then converted back into audio data and provided to the student via the terminal.

[0490] As a concrete example of a prompt, by inputting "Please describe medieval European clothing in detail," the generating AI model can extract relevant information and provide a detailed explanation. In this way, the present invention realizes improved efficiency and interactive learning support in educational settings.

[0491] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0492] Step 1:

[0493] The server retrieves digital learning materials from cloud storage and saves them to its storage device. The input includes the file format of the digital learning materials. Upon receiving a request to retrieve materials, the server searches for the relevant content and prepares the data. The output is a dataset in which the learning material content is available for access.

[0494] Step 2:

[0495] The server processes the educational material data using a text-to-speech (Speech Synthesizer) to convert it into audio data. Specifically, it utilizes speech technologies such as the Google Text-to-Speech API to convert text data into audio files. In this process, the text data is input into the Speech Synthesizer and output as audio data.

[0496] Step 3:

[0497] The server streams the generated audio data to the terminals. This ensures that audio for use in class is delivered in real time. The input is the converted audio data, and the output is the audio streaming to the terminals via the network.

[0498] Step 4:

[0499] The terminal receives audio data from the server and outputs it within the classroom via a playback device. For playback to begin, audio data must be input to the terminal. Specifically, the audio is played throughout the classroom via speakers. The output is the audio audible to the class participants.

[0500] Step 5:

[0501] The terminal receives voice input from the learner and converts it into query data using an input device. The input includes the learner's question, which is collected via a microphone and processed into digital data. The resulting output is query data that is sent to the server.

[0502] Step 6:

[0503] The server processes the query data using a response generator and generates an appropriate response using a generation AI model. This process involves designing a prompt and inputting it into the model to obtain a response in appropriate text format. The output is the response information that should be converted into audio data.

[0504] Step 7:

[0505] The server converts the generated text responses into speech and then sends them to the terminal. A text-to-speech synthesis device is used for this conversion. The input is the response generated by the AI ​​model, and the output is the audio data streamed to the terminal.

[0506] Step 8:

[0507] The terminal finally outputs the audio data received from the server back into the classroom via a playback device. This completes the feedback process to the learners. The input is audio data, and the output is an audible response.

[0508] (Application Example 1)

[0509] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0510] Current education systems struggle to provide interactive learning experiences in the classroom, particularly in real-time question-and-answer sessions where quick and accurate feedback is difficult to obtain. Furthermore, the audio guides and explanations users receive when searching for information in virtual environments are limited, leaving much room for improvement in the depth of learning. This, in turn, limits the provision of efficient learning environments and the quality of education, posing significant challenges.

[0511] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0512] In this invention, the server includes data storage means, voice generation means, and voice output means. This enables users to search for information in a virtual environment in real time and receive immediate responses from AI.

[0513] A "data storage device" is a device that can store digital learning materials and retrieve them as needed.

[0514] A "speech generation means" is a device that creates audio data based on stored digital learning materials.

[0515] "Audio output means" refers to a device for physically playing back the generated audio data and providing it to the user.

[0516] A "voice reception device" is a device that recognizes voice input from a user and converts it into digital data.

[0517] A "response generation device" is a device that constructs and responds to appropriate information in real time based on inquiries from users.

[0518] A "response output device" is a device for reproducing the generated response and communicating it to the user.

[0519] An "exploration tool" is a device that provides guidance to users when they search for and learn information in a virtual space.

[0520] An "interactive response system" is a device that uses AI to enable immediate responses to user questions.

[0521] To implement this invention, a system is configured to support learning and exploration in a virtual environment. The server stores digital learning materials and uses a speech generation means to create audio data related to the information requested by the user. The speech output means transmits this audio data to a terminal and plays it back in a format audible to the user. The terminal is equipped with a speech receiving means that receives the user's voice input, converts it into digital data, and transmits it to the server.

[0522] The server utilizes response generation mechanisms and constructs instant responses using AI algorithms. The generated responses are then converted back into audio data and delivered to the user through response output mechanisms. This allows users to gain an interactive learning experience in a virtual environment.

[0523] As a concrete example, if a user visits a virtual history museum and asks, "Who are the major artists of the Renaissance?", the server will use AI to instantly generate an answer such as, "Michelangelo and Leonardo da Vinci are representative artists of the Renaissance," and provide it in voice. This process utilizes Google's speech recognition API and a question-answering model from the transformers library.

[0524] The following are examples of input prompts for a generative AI model.

[0525] Question: Who were the major artists of the Renaissance?

[0526] Context: Information about Renaissance art and history.

[0527] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0528] Step 1:

[0529] The user enters a virtual environment and requests an audio guide on a specific topic. The terminal uses data storage from a server to retrieve relevant digital learning materials, which triggers the use of the audio generation system.

[0530] Step 2:

[0531] The server utilizes speech generation technology to create audio data from digital learning materials. The input consists of user-specified topic information and related context data. The output is audio data. A program within the server encodes text data into audio data.

[0532] Step 3:

[0533] The server transfers the generated audio data to the terminal, which then plays the audio using its audio output device. By listening to this, the user receives guide information within the virtual environment.

[0534] Step 4:

[0535] The user uses a voice input system to verbally request further information or ask specific questions to the terminal. This speech is transmitted to the terminal as voice input. Voice recognition software converts the input into text data and sends it to the server as query data.

[0536] Step 5:

[0537] The server uses an answer generation mechanism to generate immediate responses based on text data and leveraging an AI model. User inquiry data and digital learning materials serve as input. The output is a clear, textual answer to the user's question. The AI ​​model generates question-answers using contextually relevant prompts.

[0538] Step 6:

[0539] The server converts the response text back into audio data and sends it to the terminal. The terminal then plays the generated response aloud to the user using its audio output device. This allows the user to resolve any questions they may have on the spot.

[0540] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0541] This invention is a system for conducting lessons based on digital educational materials, and in particular, incorporates an emotion engine that recognizes the user's emotional state and adjusts the lesson content and responses accordingly. This provides a more interactive and personalized educational experience compared to conventional systems.

[0542] Specifically, the server stores digital educational materials as information storage and supplies them to the generation means as audio data necessary for the lesson to progress. The generation means generates appropriate audio for the lesson topic selected by the user and plays it back through the audio output means. When the audio is played back, an emotion engine analyzes the user's emotional state and adjusts the tone and speed of the audio based on the results.

[0543] The terminal conducts lessons through an audio output device and accepts user input as needed. During lessons, users communicate questions and feedback from students to the terminal, which is then converted into digital data by the input reception device and sent to the server.

[0544] The server generates an answer using a response generation mechanism based on the received question data. Furthermore, the emotion engine adjusts the tone and content of the answer according to the user's emotional state. For example, if the user appears confused, the server generates a response that includes a detailed explanation or a simple example. The generated answer is then converted back into audio data and played back through the terminal using an audio output mechanism.

[0545] This allows users to provide direct instructions and additional supplementary information, which is expected to improve the quality of lessons. By incorporating an emotion engine, the system can understand the user's emotions in real time and provide an optimal lesson experience. For example, if a student asks a difficult question and the user shows a confused expression, the emotion engine will recognize this state and the server will adjust the voice to provide a simpler and easier-to-understand explanation. This allows the user to continue the lesson efficiently.

[0546] The following describes the processing flow.

[0547] Step 1:

[0548] The server stores digital educational materials in an information storage device and makes them accessible. The server prepares to supply the materials to the generation device as needed.

[0549] Step 2:

[0550] The user operates their device to start a lesson and select a lesson topic. This information is sent from the device to the server as a request.

[0551] Step 3:

[0552] The server selects educational materials corresponding to the lesson topic and generates audio data using a generation method. The server then sends this data to the terminal.

[0553] Step 4:

[0554] The terminal plays the audio data received from the server using the audio output device, and the lesson begins. At this time, the emotion engine analyzes the user's facial expressions and tone of voice to determine their emotional state.

[0555] Step 5:

[0556] The emotion engine analyzes the user's emotional state in real time and adjusts the tone and speed of the audio during the lesson based on the information obtained. For example, if the user appears relaxed, the tone will be made calmer.

[0557] Step 6:

[0558] When a student has a question, the user receives the question using the input method on their device. This data is then converted into a digital format and sent to the server.

[0559] Step 7:

[0560] The server receives the question data and generates the optimal answer using a response generation mechanism. The server adjusts the answer based on feedback from the emotion engine; for example, if it determines that the user is confused, it will include a clearer explanation.

[0561] Step 8:

[0562] The terminal plays back the audio data of the answer received from the server using the response output device, providing the answer to the student. This allows the student to obtain a solution to their question.

[0563] Step 9:

[0564] The user monitors the lesson and supports student understanding by adjusting the audio in real time and providing additional explanations as needed. Finally, when the lesson ends, the system is shut down via the terminal.

[0565] (Example 2)

[0566] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0567] Traditional education systems often conducted lessons unilaterally without considering students' emotional states, making it difficult to provide an optimized learning experience for each individual learner. Furthermore, they were unable to immediately address students' confusion and anxiety when understanding complex material, resulting in decreased learning efficiency.

[0568] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0569] In this invention, the server includes means for storing information, processing means for receiving data from the information and generating speech, and means for outputting the speech generated by the processing means. This enables real-time analysis of the learner's emotional state and adjustment of speech tone and speed based on the results. This provides a personalized and interactive educational experience and improves the quality of learning.

[0570] "Means for storing information" refers to devices or systems that store digital educational materials and related data on media such as storage devices, and make them accessible as needed.

[0571] "A means of generating speech" refers to a device or program that creates synthesized speech from text data or other information.

[0572] "Means for outputting sound" refers to a device or interface for playing back the generated sound through speakers or headphones.

[0573] "An input means that receives voice input from participants and generates data" refers to a device or system that captures participants' speech using an input device such as a microphone, analyzes it, and converts it into digital data.

[0574] "Generating means for creating voice responses" refers to a program or system that generates appropriate response messages based on the participant's questions and feedback.

[0575] "Output means for reproducing a voice response" refers to a device or system that reproduces the generated voice response through an audio output device.

[0576] "An adjustment method that analyzes the emotional state of the participant and adjusts the tone and speed of the voice" refers to a device or system that judges the participant's emotions from their facial expressions and voice, and dynamically changes the characteristics of the generated voice based on that information.

[0577] A "correction mechanism for modifying response content based on emotions" refers to a device or program that appropriately modifies the content or expression of an existing response, taking into account the emotional state of the participant.

[0578] This invention is a lesson progression system based on digital educational materials that analyzes the emotional state of students in real time and appropriately adjusts the lesson content to provide an individualized educational experience. Specifically, the system is configured as follows.

[0579] The server stores digital educational materials using information storage means. A database is used for managing and providing educational materials. When a lesson begins, the server selects relevant materials according to the lesson topic chosen by the user and generates them as audio data using processing means. In this process, text-to-speech software or a speech synthesis engine is used to generate high-quality audio from the text.

[0580] The device plays back audio generated through its speaker via an audio output device, and also monitors the participant's facial expressions and behavior using its built-in camera and sensors. This allows the device to continuously collect and analyze data necessary for the emotion engine. If the participant shows signs of confusion or misunderstanding, the device sends instructions to the server to adjust the tone and pace of the audio. It also accepts voice input from the user, generates question data, and sends it to the server.

[0581] When a student asks a question, the server uses a response generation mechanism to return an appropriate answer to the student. Utilizing a generative AI model, it creates detailed and easy-to-understand answers based on the question, with an emotion engine flexibly adjusting the tone and content as needed. The generated answer is then converted back into audio data and provided to the student via their device.

[0582] For example, if a student asks a question about a scientific topic and shows signs of confusion, the server may generate a modified response such as, "This concept is often perceived as difficult, so let me explain it in more detail." An example of a prompt would be, "If a user shows signs of confusion during class, simplify the explanation and soften the tone."

[0583] This system, with its dynamic content adjustments tailored to each student, is expected to improve learning efficiency and maximize comprehension.

[0584] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0585] Step 1:

[0586] The server retrieves digital educational materials related to the user's selected lesson topic from its information storage device. The input is the selected lesson topic information, and the output is the corresponding lesson material data. The lesson material data is retrieved as a text file.

[0587] Step 2:

[0588] The server converts the acquired educational material data into audio data using processing tools. The input is text-based educational material data, and the output is a synthesized audio file. Specifically, text-to-speech software is used to convert the text into speech.

[0589] Step 3:

[0590] The terminal receives audio data transmitted from the server and plays it back to the learner through an audio output device. The input is an audio file, and the output is audio playback to the learner. Specifically, the terminal plays the generated audio through its speaker.

[0591] Step 4:

[0592] The device uses cameras and sensors to collect real-time facial expression data from participants. Visual data of the participants is the input, and facial expression data is obtained as the output. Specifically, facial recognition software is used to analyze the participants' emotional state.

[0593] Step 5:

[0594] The emotion engine determines the learner's emotional state based on acquired facial expression data and suggests adjustments to voice tone and speed. The input is facial expression analysis data, and the output is adjustment instructions. Based on this, the audio being played is dynamically adjusted.

[0595] Step 6:

[0596] The user enters a question into the terminal, and the input receiving device sends it to the server as voice or text data. The input is the student's question data, and the output is sent to the server in digital format. Specifically, speech recognition software is used to convert the voice into text.

[0597] Step 7:

[0598] The server generates appropriate answers using a generative AI model based on the received question data. The input is the question data, and the output is a speech-convertible answer text. Specifically, the AI ​​model presents information from a relevant knowledge base.

[0599] Step 8:

[0600] The generated response is converted into audio data, sent to the terminal, and played back by the audio output device. The input is response text data, and the output is an audio response. Specifically, a speech synthesis engine is used to generate and play back a smooth response.

[0601] This series of steps enables interactive and personalized education that responds to the emotional state of the participants.

[0602] (Application Example 2)

[0603] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0604] In modern education, a challenge is providing an educational environment that is tailored to each learner's individual level of understanding and emotional state. Traditional educational systems have struggled to customize lesson progression and responses to accommodate learners' emotions, making it difficult to provide optimal education that meets individual needs. Especially in online and virtual environments, there is a need for mechanisms that maintain learners' interest and concentration and promote effective learning.

[0605] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0606] In this invention, the server includes an information storage means for providing digital learning materials, an emotion analysis means for analyzing the learner's emotional state, and an educational content adjustment means for dynamically adjusting the educational content based on the analysis results. This makes it possible to provide interactive and personalized educational content that responds to the learner's emotional state.

[0607] "Digital learning materials" refer to educational content and teaching materials that are available on computer systems and digital devices and are stored electronically.

[0608] "Information storage means" refers to physical or electronic devices or systems for storing digital learning materials and related information.

[0609] "Generation means" refers to a mechanism or process for forming content such as audio based on data received from information storage means.

[0610] "Audio output means" refers to a device or system used to allow a user to hear the generated audio.

[0611] "Input receiving means" refers to a device or method that receives voice input from learners and processes it as necessary information.

[0612] "Question data" refers to formalized information containing questions and inquiries that learners raise during class.

[0613] "Response generation means" refers to a system or method for constructing appropriate answers or responses based on question data.

[0614] "Response output means" refers to a device or function for transmitting the generated response to the user.

[0615] "Emotional analysis methods" refer to techniques and technologies used to analyze a user's emotional state and adjust the system's operation based on that information.

[0616] "Educational content adjustment means" refers to a mechanism that appropriately adjusts the educational content provided based on the results of an analysis of emotional states.

[0617] To implement this invention, the program must be installed on both the server and the user's terminal. The server stores digital learning materials in an information storage means and generates appropriate audio based on this information using a generation means. The generated audio is output to the user's terminal via an audio output means.

[0618] The user's terminal receives voice input from the learner using an input receiving means, generates question data, and sends it to the server. The server uses a response generation means to create an answer based on the question data. At that time, an emotion analysis means analyzes the user's emotional state from facial expressions and other factors, and adjusts the tone and content of the response based on the results.

[0619] The voice response is played back by the response output device, and the educational content adjustment device dynamically modifies the learning materials according to the analyzed emotion. This makes adjustments that allow learners to understand more effectively.

[0620] For example, if a user shows a confused expression regarding an assignment they are working on, the emotion analysis tool will detect this. Based on the analysis results, the server adjusts the educational content to add simple and visually easy-to-understand examples. For instance, a prompt such as, "If a student is bored, suggest how to make learning more engaging," is used to have the AI ​​model generate content. In this way, by utilizing prompts generated by a generative AI model, a customized educational experience can be provided to the user.

[0621] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0622] Step 1:

[0623] The server retrieves digital learning materials from the information storage device. If the information is recorded as text, images, or audio data, it identifies the necessary learning materials and converts them to the appropriate format. The input is data from the information storage device, and the output is learning material data for the generation device. The server organizes the digital materials and prepares them for generation.

[0624] Step 2:

[0625] The server uses a generation mechanism to generate audio from the received educational material data. During this process, speech synthesis technology is used to create content with an easy-to-understand volume and tone, depending on the content of the material. The input is the educational material data, and the output is audio data. The server executes the speech synthesis process and prepares the audio for transmission to the terminal.

[0626] Step 3:

[0627] The terminal plays the audio data received via its audio output device. This operation is performed through the terminal's speaker, enabling the user to begin learning. The input is audio data from the server, and the output is the audio the user hears. The terminal plays the audio and provides information to the user.

[0628] Step 4:

[0629] The user communicates opinions and questions to the terminal via voice input through an input receiving device. This data is digitized by the terminal and generated as question data. The input is the user's voice, and the output is question data sent to the server. The user clearly expresses their question, and the terminal creates the necessary data.

[0630] Step 5:

[0631] The server receives question data and analyzes it using a response generation system. It generates an answer that matches the content of the question and also considers the user's emotional state using an emotion analysis system. Based on this information, the server adjusts the content of the answer. The input is question data and emotion information, and the output is the adjusted response data. The server dynamically generates answers to improve relevance.

[0632] Step 6:

[0633] The terminal uses a response output mechanism to play back the response data provided by the server as audio. At this stage, the learner receives feedback from the server. The input is the response data from the server, and the output is the audio the learner hears. The terminal converts the answer into audio, facilitating smooth communication with the user.

[0634] Step 7:

[0635] The server activates an educational content adjustment mechanism based on information obtained from sentiment analysis, and reconfigures learning materials as needed. It utilizes a generative AI model to generate new educational content using prompts, changing the next learning objective. Input is the user's sentiment data and prompts, and output is the adjusted learning content. The server continuously optimizes the educational content to support the user's learning.

[0636] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0637] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0638] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0639] [Fourth Embodiment]

[0640] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0641] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0642] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0643] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0644] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0645] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0646] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0647] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0648] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0649] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0650] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0651] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0652] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0653] This invention is a system for managing lesson progress and responding to student questions based on digitized educational materials. The following describes a specific implementation of this system.

[0654] First, at the heart of the system resides digital educational materials, and the information storage means that stores these materials acts as a server. The server accesses the digital educational materials and, based on this, supplies the audio necessary for the lesson to the audio generation means. The server provides the necessary educational content according to the data requested by the audio generation means and distributes the generated audio data to the terminals.

[0655] Next, the terminal plays the audio data received from the server and conducts the classroom lesson through the audio output device. The terminal selects appropriate audio according to the context of the lesson and supports the flow of the lesson. The terminal is also equipped with an input device to receive questions from students via voice. The questions entered via voice are converted into digital data and sent to the server.

[0656] The server generates appropriate answers from digital educational materials based on the received question data. To do this, the server uses response generation means and AI algorithms to quickly construct appropriate answers. The answers are then sent back to the terminal as audio data.

[0657] The terminal provides responses to student questions by playing back audio data received from the server using its response output mechanism. This allows for immediate feedback on student inquiries without interrupting the flow of the lesson.

[0658] The user (teacher) can monitor the progress of the lesson, intervene in the device as needed, and pause or add explanations to the audio. This allows the teacher to confirm students' understanding and take on roles such as supplementing key points.

[0659] As a concrete example, when lecturing on "Edo period culture" in a history class, the server generates audio for the relevant topic from a digitized history textbook, and the terminals play this audio in the classroom. If a student asks, "I want to know more about Edo period clothing," the terminal transmits this as voice input to the server, which uses a response generation mechanism to provide a detailed explanation and returns it to the student as voice output.

[0660] Thus, this system is designed to efficiently support lessons while reducing the burden on teachers.

[0661] The following describes the processing flow.

[0662] Step 1:

[0663] The server stores digital educational materials in an information storage device and makes them accessible as needed. The server organizes teaching material data related to topics used in lessons and prepares it so that the generation device can efficiently generate audio.

[0664] Step 2:

[0665] At the start of a class, the user operates their device and selects a class topic. This action causes the device to send a request to the server based on the selected topic and begin retrieving the necessary educational materials.

[0666] Step 3:

[0667] The server, in response to the user's request, passes educational materials on the selected topic to the audio generation device. The server then converts the specified content into audio data and prepares it for transmission to the terminal.

[0668] Step 4:

[0669] The terminal plays back audio data received from the server using an audio output device. This allows the lesson to begin with an audio guide such as, "Today we will learn about XX."

[0670] Step 5:

[0671] When a student asks a question, the terminal's input receiving mechanism receives it as voice input and converts it into digital data. The terminal then sends the converted question data to the server and requests the generation of a response.

[0672] Step 6:

[0673] The server receives the question data and uses a response generation mechanism to create an answer based on digital educational materials. The server generates the optimal answer to the question and sends it to the terminal as audio data.

[0674] Step 7:

[0675] The device plays back the received audio data of the response via a response output device, providing the student with the answer to their question. This allows students to resolve their questions without disrupting the flow of the lesson.

[0676] Step 8:

[0677] The user monitors the lesson progress and controls audio playback via their device as needed. The user can provide supplementary explanations and adjust the lesson pace based on student reactions.

[0678] Step 9:

[0679] When the lesson ends, the user exits the lesson by operating their terminal. The server creates a lesson log and stores data to help improve future lessons.

[0680] (Example 1)

[0681] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0682] In today's educational environment, teachers spend a great deal of time and effort managing the progress of lessons and responding to individual student questions. As a result, there is a risk of insufficient time for home education and a decline in the quality of lesson content. In particular, the burden on teachers increases when individualized support tailored to each student's level of understanding is required. To address these challenges, a system is needed that allows for efficient lesson progression and immediate responses to individual student questions.

[0683] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0684] In this invention, the server includes a storage device means for providing digital teaching materials, a speech synthesis device means for acquiring digital teaching materials from the storage device and generating audio, and an input device means for receiving voice input from learners and generating inquiry data. This makes it possible to respond to students' questions in real time while smoothly conducting lessons.

[0685] A "storage device" is a device used to store and keep digital learning materials accessible.

[0686] A "speech synthesis device" is a device that generates audio data from information acquired from digital educational materials.

[0687] A "playback device" is a device that outputs generated audio data and allows learners to listen to it.

[0688] An "input device" is a device that receives voice input from learners and processes it as inquiry data.

[0689] A "response generation device" is a device that generates an appropriate voice response based on inquiry data.

[0690] An "output device" is a device that plays back the generated voice response and provides it to the learner.

[0691] A "generative artificial intelligence model" is a computational model that utilizes information technology to generate optimized responses based on query data.

[0692] This invention relates to an educational support system that utilizes digital teaching materials, and is configured in which a server, a terminal, and a user work together in a coordinated manner.

[0693] The server is the core of the educational support system. It stores digital learning materials in cloud storage and functions as a memory device. When a request for learning materials is received, the server uses a speech synthesis device and utilizes speech technologies such as the Google Text-to-Speech API to convert the materials into audio data. The converted audio data is then streamed to the terminal.

[0694] The terminal is a device that plays audio data received from the server. Equipped with a playback device, the terminal supports the progress of lessons by outputting audio within the classroom. The terminal also has an input device that receives audio input from learners and transmits it to the server. The input questions are processed appropriately and converted into digital data.

[0695] The user (teacher), acting as the system's supervisor, pauses audio output and provides additional explanations via their terminal as needed during lessons. This allows for flexible lesson progression based on the learners' understanding. The user also reviews learners' questions and provides supplementary information where further explanation is required.

[0696] As a concrete example, consider a scenario in a history class where a lecture is given on "Medieval European Culture." The server generates audio for the relevant topic from digitized history materials, and the terminals play this audio in the classroom. For example, if a student asks, "Please tell me about medieval European clothing," the terminal uses speech recognition technology to transmit the question to the server. The server uses a response generation device and a generative AI model to generate an appropriate answer. This answer is then converted back into audio data and provided to the student via the terminal.

[0697] As a concrete example of a prompt, by inputting "Please describe medieval European clothing in detail," the generating AI model can extract relevant information and provide a detailed explanation. In this way, the present invention realizes improved efficiency and interactive learning support in educational settings.

[0698] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0699] Step 1:

[0700] The server retrieves digital learning materials from cloud storage and saves them to its storage device. The input includes the file format of the digital learning materials. Upon receiving a request to retrieve materials, the server searches for the relevant content and prepares the data. The output is a dataset in which the learning material content is available for access.

[0701] Step 2:

[0702] The server processes the educational material data using a text-to-speech (Speech Synthesizer) to convert it into audio data. Specifically, it utilizes speech technologies such as the Google Text-to-Speech API to convert text data into audio files. In this process, the text data is input into the Speech Synthesizer and output as audio data.

[0703] Step 3:

[0704] The server streams the generated audio data to the terminals. This ensures that audio for use in class is delivered in real time. The input is the converted audio data, and the output is the audio streaming to the terminals via the network.

[0705] Step 4:

[0706] The terminal receives audio data from the server and outputs it within the classroom via a playback device. For playback to begin, audio data must be input to the terminal. Specifically, the audio is played throughout the classroom via speakers. The output is the audio audible to the class participants.

[0707] Step 5:

[0708] The terminal receives voice input from the learner and converts it into query data using an input device. The input includes the learner's question, which is collected via a microphone and processed into digital data. The resulting output is query data that is sent to the server.

[0709] Step 6:

[0710] The server processes the query data using a response generator and generates an appropriate response using a generation AI model. This process involves designing a prompt and inputting it into the model to obtain a response in appropriate text format. The output is the response information that should be converted into audio data.

[0711] Step 7:

[0712] The server converts the generated text responses into speech and then sends them to the terminal. A text-to-speech synthesis device is used for this conversion. The input is the response generated by the AI ​​model, and the output is the audio data streamed to the terminal.

[0713] Step 8:

[0714] The terminal finally outputs the audio data received from the server back into the classroom via a playback device. This completes the feedback process to the learners. The input is audio data, and the output is an audible response.

[0715] (Application Example 1)

[0716] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0717] Current education systems struggle to provide interactive learning experiences in the classroom, particularly in real-time question-and-answer sessions where quick and accurate feedback is difficult to obtain. Furthermore, the audio guides and explanations users receive when searching for information in virtual environments are limited, leaving much room for improvement in the depth of learning. This, in turn, limits the provision of efficient learning environments and the quality of education, posing significant challenges.

[0718] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0719] In this invention, the server includes data storage means, voice generation means, and voice output means. This enables users to search for information in a virtual environment in real time and receive immediate responses from AI.

[0720] A "data storage device" is a device that can store digital learning materials and retrieve them as needed.

[0721] A "speech generation means" is a device that creates audio data based on stored digital learning materials.

[0722] "Audio output means" refers to a device for physically playing back the generated audio data and providing it to the user.

[0723] A "voice reception device" is a device that recognizes voice input from a user and converts it into digital data.

[0724] A "response generation device" is a device that constructs and responds to appropriate information in real time based on inquiries from users.

[0725] A "response output device" is a device for reproducing the generated response and communicating it to the user.

[0726] An "exploration tool" is a device that provides guidance to users when they search for and learn information in a virtual space.

[0727] An "interactive response system" is a device that uses AI to enable immediate responses to user questions.

[0728] To implement this invention, a system is configured to support learning and exploration in a virtual environment. The server stores digital learning materials and uses a speech generation means to create audio data related to the information requested by the user. The speech output means transmits this audio data to a terminal and plays it back in a format audible to the user. The terminal is equipped with a speech receiving means that receives the user's voice input, converts it into digital data, and transmits it to the server.

[0729] The server utilizes response generation mechanisms and constructs instant responses using AI algorithms. The generated responses are then converted back into audio data and delivered to the user through response output mechanisms. This allows users to gain an interactive learning experience in a virtual environment.

[0730] As a concrete example, if a user visits a virtual history museum and asks, "Who are the major artists of the Renaissance?", the server will use AI to instantly generate an answer such as, "Michelangelo and Leonardo da Vinci are representative artists of the Renaissance," and provide it in voice. This process utilizes Google's speech recognition API and a question-answering model from the transformers library.

[0731] The following are examples of input prompts for a generative AI model.

[0732] Question: Who were the major artists of the Renaissance?

[0733] Context: Information about Renaissance art and history.

[0734] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0735] Step 1:

[0736] The user enters a virtual environment and requests an audio guide on a specific topic. The terminal uses data storage from a server to retrieve relevant digital learning materials, which triggers the use of the audio generation system.

[0737] Step 2:

[0738] The server utilizes speech generation technology to create audio data from digital learning materials. The input consists of user-specified topic information and related context data. The output is audio data. A program within the server encodes text data into audio data.

[0739] Step 3:

[0740] The server transfers the generated audio data to the terminal, which then plays the audio using its audio output device. By listening to this, the user receives guide information within the virtual environment.

[0741] Step 4:

[0742] The user uses a voice input system to verbally request further information or ask specific questions to the terminal. This speech is transmitted to the terminal as voice input. Voice recognition software converts the input into text data and sends it to the server as query data.

[0743] Step 5:

[0744] The server uses an answer generation mechanism to generate immediate responses based on text data and leveraging an AI model. User inquiry data and digital learning materials serve as input. The output is a clear, textual answer to the user's question. The AI ​​model generates question-answers using contextually relevant prompts.

[0745] Step 6:

[0746] The server converts the response text back into audio data and sends it to the terminal. The terminal then plays the generated response aloud to the user using its audio output device. This allows the user to resolve any questions they may have on the spot.

[0747] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0748] This invention is a system for conducting lessons based on digital educational materials, and in particular, incorporates an emotion engine that recognizes the user's emotional state and adjusts the lesson content and responses accordingly. This provides a more interactive and personalized educational experience compared to conventional systems.

[0749] Specifically, the server stores digital educational materials as information storage and supplies them to the generation means as audio data necessary for the lesson to progress. The generation means generates appropriate audio for the lesson topic selected by the user and plays it back through the audio output means. When the audio is played back, an emotion engine analyzes the user's emotional state and adjusts the tone and speed of the audio based on the results.

[0750] The terminal conducts lessons through an audio output device and accepts user input as needed. During lessons, users communicate questions and feedback from students to the terminal, which is then converted into digital data by the input reception device and sent to the server.

[0751] The server generates an answer using a response generation mechanism based on the received question data. Furthermore, the emotion engine adjusts the tone and content of the answer according to the user's emotional state. For example, if the user appears confused, the server generates a response that includes a detailed explanation or a simple example. The generated answer is then converted back into audio data and played back through the terminal using an audio output mechanism.

[0752] This allows users to provide direct instructions and additional supplementary information, which is expected to improve the quality of lessons. By incorporating an emotion engine, the system can understand the user's emotions in real time and provide an optimal lesson experience. For example, if a student asks a difficult question and the user shows a confused expression, the emotion engine will recognize this state and the server will adjust the voice to provide a simpler and easier-to-understand explanation. This allows the user to continue the lesson efficiently.

[0753] The following describes the processing flow.

[0754] Step 1:

[0755] The server stores digital educational materials in an information storage device and makes them accessible. The server prepares to supply the materials to the generation device as needed.

[0756] Step 2:

[0757] The user operates their device to start a lesson and select a lesson topic. This information is sent from the device to the server as a request.

[0758] Step 3:

[0759] The server selects educational materials corresponding to the lesson topic and generates audio data using a generation method. The server then sends this data to the terminal.

[0760] Step 4:

[0761] The terminal plays the audio data received from the server using the audio output device, and the lesson begins. At this time, the emotion engine analyzes the user's facial expressions and tone of voice to determine their emotional state.

[0762] Step 5:

[0763] The emotion engine analyzes the user's emotional state in real time and adjusts the tone and speed of the audio during the lesson based on the information obtained. For example, if the user appears relaxed, the tone will be made calmer.

[0764] Step 6:

[0765] When a student has a question, the user receives the question using the input method on their device. This data is then converted into a digital format and sent to the server.

[0766] Step 7:

[0767] The server receives the question data and generates the optimal answer using a response generation mechanism. The server adjusts the answer based on feedback from the emotion engine; for example, if it determines that the user is confused, it will include a clearer explanation.

[0768] Step 8:

[0769] The terminal plays back the audio data of the answer received from the server using the response output device, providing the answer to the student. This allows the student to obtain a solution to their question.

[0770] Step 9:

[0771] The user monitors the lesson and supports student understanding by adjusting the audio in real time and providing additional explanations as needed. Finally, when the lesson ends, the system is shut down via the terminal.

[0772] (Example 2)

[0773] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0774] Traditional education systems often conducted lessons unilaterally without considering students' emotional states, making it difficult to provide an optimized learning experience for each individual learner. Furthermore, they were unable to immediately address students' confusion and anxiety when understanding complex material, resulting in decreased learning efficiency.

[0775] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0776] In this invention, the server includes means for storing information, processing means for receiving data from the information and generating speech, and means for outputting the speech generated by the processing means. This enables real-time analysis of the learner's emotional state and adjustment of speech tone and speed based on the results. This provides a personalized and interactive educational experience and improves the quality of learning.

[0777] "Means for storing information" refers to devices or systems that store digital educational materials and related data on media such as storage devices, and make them accessible as needed.

[0778] "A means of generating speech" refers to a device or program that creates synthesized speech from text data or other information.

[0779] "Means for outputting sound" refers to a device or interface for playing back the generated sound through speakers or headphones.

[0780] "An input means that receives voice input from participants and generates data" refers to a device or system that captures participants' speech using an input device such as a microphone, analyzes it, and converts it into digital data.

[0781] "Generating means for creating voice responses" refers to a program or system that generates appropriate response messages based on the participant's questions and feedback.

[0782] "Output means for reproducing a voice response" refers to a device or system that reproduces the generated voice response through an audio output device.

[0783] "An adjustment method that analyzes the emotional state of the participant and adjusts the tone and speed of the voice" refers to a device or system that judges the participant's emotions from their facial expressions and voice, and dynamically changes the characteristics of the generated voice based on that information.

[0784] A "correction mechanism for modifying response content based on emotions" refers to a device or program that appropriately modifies the content or expression of an existing response, taking into account the emotional state of the participant.

[0785] This invention is a lesson progression system based on digital educational materials that analyzes the emotional state of students in real time and appropriately adjusts the lesson content to provide an individualized educational experience. Specifically, the system is configured as follows.

[0786] The server stores digital educational materials using information storage means. A database is used for managing and providing educational materials. When a lesson begins, the server selects relevant materials according to the lesson topic chosen by the user and generates them as audio data using processing means. In this process, text-to-speech software or a speech synthesis engine is used to generate high-quality audio from the text.

[0787] The device plays back audio generated through its speaker via an audio output device, and also monitors the participant's facial expressions and behavior using its built-in camera and sensors. This allows the device to continuously collect and analyze data necessary for the emotion engine. If the participant shows signs of confusion or misunderstanding, the device sends instructions to the server to adjust the tone and pace of the audio. It also accepts voice input from the user, generates question data, and sends it to the server.

[0788] When a student asks a question, the server uses a response generation mechanism to return an appropriate answer to the student. Utilizing a generative AI model, it creates detailed and easy-to-understand answers based on the question, with an emotion engine flexibly adjusting the tone and content as needed. The generated answer is then converted back into audio data and provided to the student via their device.

[0789] For example, if a student asks a question about a scientific topic and shows signs of confusion, the server may generate a modified response such as, "This concept is often perceived as difficult, so let me explain it in more detail." An example of a prompt would be, "If a user shows signs of confusion during class, simplify the explanation and soften the tone."

[0790] This system, with its dynamic content adjustments tailored to each student, is expected to improve learning efficiency and maximize comprehension.

[0791] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0792] Step 1:

[0793] The server retrieves digital educational materials related to the user's selected lesson topic from its information storage device. The input is the selected lesson topic information, and the output is the corresponding lesson material data. The lesson material data is retrieved as a text file.

[0794] Step 2:

[0795] The server converts the acquired educational material data into audio data using processing tools. The input is text-based educational material data, and the output is a synthesized audio file. Specifically, text-to-speech software is used to convert the text into speech.

[0796] Step 3:

[0797] The terminal receives audio data transmitted from the server and plays it back to the learner through an audio output device. The input is an audio file, and the output is audio playback to the learner. Specifically, the terminal plays the generated audio through its speaker.

[0798] Step 4:

[0799] The device uses cameras and sensors to collect real-time facial expression data from participants. Visual data of the participants is the input, and facial expression data is obtained as the output. Specifically, facial recognition software is used to analyze the participants' emotional state.

[0800] Step 5:

[0801] The emotion engine determines the learner's emotional state based on acquired facial expression data and suggests adjustments to voice tone and speed. The input is facial expression analysis data, and the output is adjustment instructions. Based on this, the audio being played is dynamically adjusted.

[0802] Step 6:

[0803] The user enters a question into the terminal, and the input receiving device sends it to the server as voice or text data. The input is the student's question data, and the output is sent to the server in digital format. Specifically, speech recognition software is used to convert the voice into text.

[0804] Step 7:

[0805] The server generates appropriate answers using a generative AI model based on the received question data. The input is the question data, and the output is a speech-convertible answer text. Specifically, the AI ​​model presents information from a relevant knowledge base.

[0806] Step 8:

[0807] The generated response is converted into audio data, sent to the terminal, and played back by the audio output device. The input is response text data, and the output is an audio response. Specifically, a speech synthesis engine is used to generate and play back a smooth response.

[0808] This series of steps enables interactive and personalized education that responds to the emotional state of the participants.

[0809] (Application Example 2)

[0810] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0811] In modern education, a challenge is providing an educational environment that is tailored to each learner's individual level of understanding and emotional state. Traditional educational systems have struggled to customize lesson progression and responses to accommodate learners' emotions, making it difficult to provide optimal education that meets individual needs. Especially in online and virtual environments, there is a need for mechanisms that maintain learners' interest and concentration and promote effective learning.

[0812] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0813] In this invention, the server includes an information storage means for providing digital learning materials, an emotion analysis means for analyzing the learner's emotional state, and an educational content adjustment means for dynamically adjusting the educational content based on the analysis results. This makes it possible to provide interactive and personalized educational content that responds to the learner's emotional state.

[0814] "Digital learning materials" refer to educational content and teaching materials that are available on computer systems and digital devices and are stored electronically.

[0815] "Information storage means" refers to physical or electronic devices or systems for storing digital learning materials and related information.

[0816] "Generation means" refers to a mechanism or process for forming content such as audio based on data received from information storage means.

[0817] "Audio output means" refers to a device or system used to allow a user to hear the generated audio.

[0818] "Input receiving means" refers to a device or method that receives voice input from learners and processes it as necessary information.

[0819] "Question data" refers to formalized information containing questions and inquiries that learners raise during class.

[0820] "Response generation means" refers to a system or method for constructing appropriate answers or responses based on question data.

[0821] "Response output means" refers to a device or function for transmitting the generated response to the user.

[0822] "Emotional analysis methods" refer to techniques and technologies used to analyze a user's emotional state and adjust the system's operation based on that information.

[0823] "Educational content adjustment means" refers to a mechanism that appropriately adjusts the educational content provided based on the results of an analysis of emotional states.

[0824] To implement this invention, the program must be installed on both the server and the user's terminal. The server stores digital learning materials in an information storage means and generates appropriate audio based on this information using a generation means. The generated audio is output to the user's terminal via an audio output means.

[0825] The user's terminal receives voice input from the learner using an input receiving means, generates question data, and sends it to the server. The server uses a response generation means to create an answer based on the question data. At that time, an emotion analysis means analyzes the user's emotional state from facial expressions and other factors, and adjusts the tone and content of the response based on the results.

[0826] The voice response is played back by the response output device, and the educational content adjustment device dynamically modifies the learning materials according to the analyzed emotion. This makes adjustments that allow learners to understand more effectively.

[0827] For example, if a user shows a confused expression regarding an assignment they are working on, the emotion analysis tool will detect this. Based on the analysis results, the server adjusts the educational content to add simple and visually easy-to-understand examples. For instance, a prompt such as, "If a student is bored, suggest how to make learning more engaging," is used to have the AI ​​model generate content. In this way, by utilizing prompts generated by a generative AI model, a customized educational experience can be provided to the user.

[0828] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0829] Step 1:

[0830] The server retrieves digital learning materials from the information storage device. If the information is recorded as text, images, or audio data, it identifies the necessary learning materials and converts them to the appropriate format. The input is data from the information storage device, and the output is learning material data for the generation device. The server organizes the digital materials and prepares them for generation.

[0831] Step 2:

[0832] The server uses a generation mechanism to generate audio from the received educational material data. During this process, speech synthesis technology is used to create content with an easy-to-understand volume and tone, depending on the content of the material. The input is the educational material data, and the output is audio data. The server executes the speech synthesis process and prepares the audio for transmission to the terminal.

[0833] Step 3:

[0834] The terminal plays the audio data received via its audio output device. This operation is performed through the terminal's speaker, enabling the user to begin learning. The input is audio data from the server, and the output is the audio the user hears. The terminal plays the audio and provides information to the user.

[0835] Step 4:

[0836] The user communicates opinions and questions to the terminal via voice input through an input receiving device. This data is digitized by the terminal and generated as question data. The input is the user's voice, and the output is question data sent to the server. The user clearly expresses their question, and the terminal creates the necessary data.

[0837] Step 5:

[0838] The server receives question data and analyzes it using a response generation system. It generates an answer that matches the content of the question and also considers the user's emotional state using an emotion analysis system. Based on this information, the server adjusts the content of the answer. The input is question data and emotion information, and the output is the adjusted response data. The server dynamically generates answers to improve relevance.

[0839] Step 6:

[0840] The terminal uses a response output mechanism to play back the response data provided by the server as audio. At this stage, the learner receives feedback from the server. The input is the response data from the server, and the output is the audio the learner hears. The terminal converts the answer into audio, facilitating smooth communication with the user.

[0841] Step 7:

[0842] The server activates an educational content adjustment mechanism based on information obtained from sentiment analysis, and reconfigures learning materials as needed. It utilizes a generative AI model to generate new educational content using prompts, changing the next learning objective. Input is the user's sentiment data and prompts, and output is the adjusted learning content. The server continuously optimizes the educational content to support the user's learning.

[0843] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0844] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0845] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0846] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0847] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0848] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0849] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0850] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0851] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0852] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0853] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0854] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0855] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0856] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0857] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0858] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0859] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0860] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0861] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0862] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0863] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0864] The following is further disclosed regarding the embodiments described above.

[0865] (Claim 1)

[0866] Information storage means for providing digital educational materials,

[0867] A generation means that receives digital educational materials from the aforementioned information storage means and generates audio,

[0868] Audio output means for outputting audio generated by the generation means,

[0869] An input receiving means that receives student voice input and generates question data,

[0870] A response generation means that generates an audio response corresponding to the aforementioned question data,

[0871] A response output means for reproducing the aforementioned voice response,

[0872] A system that includes this.

[0873] (Claim 2)

[0874] The system according to claim 1, wherein the generation means automatically generates audio according to the progress of the lesson.

[0875] (Claim 3)

[0876] The system according to claim 1, wherein the response generation means generates answers to question data based on digital educational materials.

[0877] "Example 1"

[0878] (Claim 1)

[0879] A storage device that provides digital teaching materials,

[0880] A speech synthesis device that acquires digital teaching materials from the aforementioned storage device and generates audio,

[0881] A playback device that outputs sound generated by the aforementioned speech synthesis device,

[0882] An input device that receives voice input from learners and generates inquiry data,

[0883] A response generation device that generates a voice response corresponding to the aforementioned inquiry data,

[0884] An output device for reproducing the aforementioned voice response,

[0885] A processing device that uses information technology to provide optimized voice responses based on query data using a generative artificial intelligence model,

[0886] A system that includes this.

[0887] (Claim 2)

[0888] The system according to claim 1, wherein the speech synthesis device automatically generates speech according to the progress of learning.

[0889] (Claim 3)

[0890] The system according to claim 1, wherein the response generation device generates a response to inquiry data based on digital teaching materials.

[0891] "Application Example 1"

[0892] (Claim 1)

[0893] A data storage means for providing digital learning materials,

[0894] A voice generation means that receives digital learning materials from the data storage means and generates audio,

[0895] A sound output means that outputs the sound generated by the sound generation means,

[0896] A voice reception means that receives voice input from users and generates inquiry data,

[0897] A response generation means that generates a voice response corresponding to the aforementioned inquiry data,

[0898] A response output means for playing back the aforementioned voice response,

[0899] An exploration method that provides audio guidance when users search for information in a virtual environment,

[0900] An interactive response means that generates an immediate response using an AI algorithm based on the questions from the aforementioned exploration means,

[0901] A system that includes this.

[0902] (Claim 2)

[0903] The system according to claim 1, wherein the voice generation means automatically generates voice according to the progress of learning, and is based on the context of the virtual environment.

[0904] (Claim 3)

[0905] The system according to claim 1, wherein the response generation means generates a response to inquiry data based on digital learning materials and provides it in voice within a virtual setting.

[0906] "Example 2 of combining an emotion engine"

[0907] (Claim 1)

[0908] Means of accumulating information,

[0909] A processing means that receives data from the aforementioned information and generates sound,

[0910] Means for outputting the sound generated by the processing means,

[0911] An input means that receives voice input from participants and generates data,

[0912] A generation means for creating an audio response corresponding to the aforementioned data,

[0913] Output means for reproducing the aforementioned voice response,

[0914] A means of adjusting the tone and speed of the voice based on the emotional state of the participants,

[0915] A correction method that modifies the content of responses based on the participants' emotions,

[0916] A system that includes this.

[0917] (Claim 2)

[0918] The system according to claim 1, wherein the processing means automatically generates voice according to the progress and adjusts the voice characteristics based on the user's emotions.

[0919] (Claim 3)

[0920] The system according to claim 1, wherein the generation means generates a response to data based on the accumulated information and modifies the content according to the user's emotions.

[0921] "Application example 2 when combining with an emotional engine"

[0922] (Claim 1)

[0923] Information storage means for providing digital learning materials,

[0924] A generation means that receives digital learning materials from the information storage means and generates audio,

[0925] Audio output means for outputting audio generated by the generation means,

[0926] An input receiving means that receives voice input from learners and generates question data,

[0927] A response generation means that generates an audio response corresponding to the aforementioned question data,

[0928] A response output means for reproducing the aforementioned voice response,

[0929] An emotion analysis method that analyzes the user's emotional state and adjusts the voice accordingly,

[0930] An educational content adjustment means that dynamically adjusts the educational content based on the analysis results,

[0931] A system that includes this.

[0932] (Claim 2)

[0933] The system according to claim 1, wherein the generation means automatically generates audio according to the progress of the lesson and the emotional state of the user.

[0934] (Claim 3)

[0935] The system according to claim 1, wherein the response generation means generates answers to question data based on digital learning materials and adjusts the tone and content of the answers according to the user's emotions. [Explanation of symbols]

[0936] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. Information storage means for providing digital educational materials, A generation means that receives digital educational materials from the aforementioned information storage means and generates audio, Audio output means for outputting audio generated by the generation means, An input receiving means that receives student voice input and generates question data, A response generation means that generates an audio response corresponding to the aforementioned question data, A response output means for reproducing the aforementioned voice response, A system that includes this.

2. The system according to claim 1, wherein the generation means automatically generates audio according to the progress of the lesson.

3. The system according to claim 1, wherein the response generation means generates answers to question data based on digital educational materials.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A