system

The system addresses inefficiencies in conventional learning by using voice input for personalized quiz selection and real-time feedback, enhancing learner efficiency and motivation through flexible, emotionally tailored education.

JP2026070278APending Publication Date: 2026-04-27SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-15
Publication Date
2026-04-27

AI Technical Summary

Technical Problem

Conventional learning methods lack flexibility, fail to provide accurate feedback due to pronunciation and intonation issues, and are not suitable for on-the-go learning, leading to inefficient and unmotivated learning experiences.

Method used

A system utilizing voice input for accessing learning content, personalized quiz selection based on learner history and skill level, and real-time feedback through speech recognition, enabling interactive and personalized learning.

Benefits of technology

Enhances learner efficiency and motivation by providing accurate, real-time feedback and allowing learning on-the-go, tailored to individual needs and emotional states.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026070278000001_ABST
    Figure 2026070278000001_ABST
Patent Text Reader

Abstract

We provide a system that can offer a more flexible and effective learning environment. [Solution] A system including a voice interface as a means for a user to request access to learning content through voice input, comprising: means for a user to request access to learning content through voice input; means for analyzing the voice input and dynamically selecting appropriate quiz questions based on the learner's learning history and skill level; means for presenting the selected quiz questions to the user via an audio output device; means for receiving the user's oral answers and converting them into corresponding text data using speech recognition technology; means for evaluating the text data to determine correctness and generate immediate feedback; means for providing the generated feedback to the user audibly; and means for recording the user's learning progress and analyzing the data to improve future learning plans.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Modern examinees and learners are required to learn efficiently within limited time. However, in conventional learning methods, text-based teaching materials are often used, and there are few opportunities for audio output, making it difficult to accurately grasp the learners' understanding. Also, it lacks convenience and is not suitable for use during movement or work. Furthermore, there is a risk of receiving incorrect feedback due to inaccurate understanding of the pronunciation and intonation of a specific language. The present invention aims to solve these problems and provide a more flexible and effective learning environment.

Means for Solving the Problems

[0005] This invention provides a system that includes a voice interface as a means for users to request access to learning content through voice input. Personalized learning is possible by using means to analyze voice input and dynamically select appropriate quiz questions based on the learner's learning history and skill level. By presenting the selected quiz questions to the user via a voice output device, the user can continue learning even while on the go or performing other tasks. The system also includes means to receive the user's verbal responses and convert them into corresponding text data using speech recognition technology, enabling accurate evaluation. Furthermore, it includes means to immediately provide the generated feedback to the user, enhance the effectiveness of learning, record progress, and analyze data to improve future learning plans. In this way, interactive learning through speech recognition and personalized feedback is realized, improving learner efficiency and effectiveness.

[0006] A "user" is a learner who uses the system to access learning content and answer quizzes.

[0007] "Voice input" refers to the verbal utterances used by users to give instructions or responses to a system.

[0008] "Learning content" refers to educational materials that include quiz questions and related information for preparing for certification exams.

[0009] "Voice interface" is a general term for speech recognition and speech output technologies that allow users to interact with systems through their voice.

[0010] "Speech recognition technology" is a technology that converts a user's speech into digital signals and analyzes them as text data.

[0011] A "quiz question" is a question posed to assess a user's knowledge, and it can take various forms, such as multiple-choice questions or written response questions.

[0012] "Dynamic selection" means selecting appropriate learning content in real time based on the user's characteristics and learning history.

[0013] "Feedback" refers to information that includes evaluations and explanations of the results provided by the system in response to the user's answers.

[0014] "Recording progress" means accumulating user learning history and achievements, and saving data for use in analysis.

[0015] A "voice output device" is a device that presents selected quiz questions and feedback to the user as audio. [Brief explanation of the drawing]

[0016] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined.

Mode for Carrying Out the Invention

[0017] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] <​​​​​​​​In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0022] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0024] [First Embodiment]

[0025] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0026] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0029] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0032] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0036] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0037] This system provides a voice interface that enables users to learn using voice and offers a personalized learning experience. The following describes how the server, terminal, and user interact and how the system operates.

[0038] The server consists of a core database and an AI speech recognition engine. The server first references the user's learning history and skill level, and then selects appropriate quiz questions from the database. In this selection process, the server automatically chooses the most suitable questions based on the user's unique learning patterns and past performance. Furthermore, upon receiving the user's voice response, the server converts it into text data via the speech recognition engine and performs a correct / incorrect judgment. Based on the judgment result, it generates feedback for the user and sends it to their terminal.

[0039] The terminal includes a voice output device and a voice input device that directly interact with the user. It converts the user's voice input into a digital signal and transmits it to the server. It also provides the user with received quiz questions via voice and prompts them to answer. The terminal assists in converting voice to text using voice recognition technology, providing data for the server to perform accurate evaluations. Furthermore, it plays back feedback sent from the server via voice, conveying the user's learning progress.

[0040] Users participate in the quiz through interaction with their device. For example, if a user answers the question "What is the capital of Japan?" with "Tokyo," the server analyzes the audio and immediately provides feedback that the answer is correct. This method allows users to understand the accuracy of their learning in real time and move on to the next question. Users can also check their progress and past performance through the app on their device or the web interface, and adjust their learning direction accordingly.

[0041] In this way, this system utilizes speech recognition technology and digital communication technology to provide a highly personalized learning experience. Users can learn at their own pace, and effective learning is possible even while on the go or working.

[0042] The following describes the processing flow.

[0043] Step 1:

[0044] The user gives a voice command to the device to begin learning. For example, they might say, "Start the quiz."

[0045] Step 2:

[0046] The terminal converts the user's voice commands into digital signals and sends them to the server as requests.

[0047] Step 3:

[0048] The server parses the received request and references the user's learning history database for past performance and learning topics.

[0049] Step 4:

[0050] The server selects appropriate quiz questions from the database based on the learning history and current skill level.

[0051] Step 5:

[0052] The server sends the selected quiz questions to the terminal and instructs it to present them to the user.

[0053] Step 6:

[0054] The device uses speech synthesis technology to convert the received quiz questions into speech and presents them to the user.

[0055] Step 7:

[0056] The user answers quiz questions by voice, and the device receives their utterance.

[0057] Step 8:

[0058] The terminal transmits the user's voice response as a digital signal to the server.

[0059] Step 9:

[0060] The server uses speech recognition technology to convert the audio signal into text data and compares it to correct answer data to evaluate the response.

[0061] Step 10:

[0062] The server performs a correct / incorrect judgment, generates a feedback message, and sends it to the terminal.

[0063] Step 11:

[0064] The device provides feedback to the user via voice. For example, it might say, "That's correct!" or "That's incorrect, the correct answer is Tokyo."

[0065] Step 12:

[0066] The device records the user's answers and sends the data to the server to be used in selecting the next quiz.

[0067] In this way, the series of processes is completed, and the user can proceed to the next quiz.

[0068] (Example 1)

[0069] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0070] Modern learning systems often provide uniform content to all learners simultaneously, lacking individualization that takes into account each learner's skill level and learning pattern. This makes effective learning difficult and hinders the ability to maintain learner motivation.

[0071] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0072] In this invention, the server includes means for the user to request access to learning resources through voice input, means for analyzing the voice input and dynamically selecting appropriate problems based on the learner's history and skill level, and means for recording the user's learning progress and analyzing the information to improve future plans. This provides a personalized learning experience tailored to the individual needs of each learner, enabling effective and highly motivating learning.

[0073] "Learning resources" refer to information or learning materials that users use to advance their learning.

[0074] "Voice input" refers to the process in which a user provides information verbally and a system receives it.

[0075] "Learner history" refers to a record showing what learning activities a user has engaged in in the past.

[0076] "Skill level" refers to an indicator that shows the current level of a user's knowledge and abilities.

[0077] "Dynamic selection" refers to the process of choosing the most appropriate option in real time or based on the situation.

[0078] An "audio output device" refers to a device that generates audio data from an electronic device in a format that the user can listen to.

[0079] "Verbal response" refers to the act of a user providing an answer via voice.

[0080] "Recognition technology" refers to technologies for converting audio and other data into text or other formats.

[0081] "The relevant text data" refers to text data converted from voice input.

[0082] "Correct / incorrect judgment" refers to the process of evaluating whether user input is correct or incorrect.

[0083] "Feedback" refers to responses or guidance provided to users regarding their actions.

[0084] "Learning progress" refers to the state indicating how much progress a user has made through their learning process.

[0085] "Personalized experience" refers to a learning experience tailored to the individual needs and characteristics of each user.

[0086] In embodiments of the present invention, the system consists of a server, a terminal, and a user, each playing a specific role.

[0087] The server is the core information processing device of this invention. The server is equipped with a speech recognition engine and a database, and receives and processes user voice data internally. A common cloud-based speech recognition service can be used as the speech recognition engine; for example, Google® Cloud Speech-to-Text API can be used. The server first analyzes the voice input and selects the most suitable question by referencing the learner's history and skill level from the database. A generative AI model is used for this selection. The selected question is then sent from the server to the terminal.

[0088] A terminal is an information device that provides an interface with the user. The terminal is equipped with a voice input device and a voice output device, and transmits learning requests entered by the user via voice to the server. In addition, it plays back questions sent from the server via voice and presents them to the user. Furthermore, it records the user's answers and sends them back to the server. Specific forms of terminals include smartphones and dedicated tablets.

[0089] Users answer questions verbally via their devices, allowing them to learn while experiencing realistic interaction. For example, if a user receives the question "What is the capital of Japan?" and answers "Tokyo" verbally, the server analyzes the answer and immediately provides feedback on whether it is correct or incorrect.

[0090] As a concrete example, a prompt message for a generative AI model could be: "Based on the user's learning history and skill level, select the next question. Then, analyze the user's voice response and generate appropriate feedback."

[0091] In this way, by having the server, terminal, and user work together, it becomes possible to provide a real-time and personalized learning experience.

[0092] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0093] Step 1:

[0094] The server receives a voice access request from the user to learning content. It takes voice data from the device as input. To analyze this voice data, it uses a speech recognition engine to perform recognition processing and convert it into text data. This process involves receiving the voice signal as digital data and converting it into text format via a speech recognition API.

[0095] Step 2:

[0096] The server analyzes the converted text data and retrieves the user's learning history and skill level from the database. The input is audio data in text format. Based on this data, the server dynamically selects appropriate learning questions using a generative AI model. It performs database queries and considers past performance and learning patterns to determine the most suitable quiz.

[0097] Step 3:

[0098] The server sends the selected problem as text data to the terminal. The input is the data of the training problems selected by the AI ​​model. The output is the data that the terminal uses speech synthesis technology to provide to the user with the selected problem in voice. This involves sending the data via the HTTP protocol and the synthesis process to convert it into speech.

[0099] Step 4:

[0100] The user responds to questions presented by the terminal using voice. The input is the user's own voice response. The terminal records this voice, converts it into a digital signal, and sends it back to the server. This is where the microphone-based voice capture and digital conversion functions come into play.

[0101] Step 5:

[0102] The server converts the voice response data received from the terminal back into text data using a speech recognition engine. The input is digitized voice data. The server analyzes this data and determines whether it is correct or incorrect. The generated text data is then passed to an evaluation algorithm, which includes an analysis process to determine whether the quiz answers are correct or incorrect.

[0103] Step 6:

[0104] The server generates a feedback message based on the judgment result and sends it to the terminal. The input is the correct / incorrect judgment result. The output is the feedback content to be provided to the user in voice. The server then converts this feedback into text and sends it to the terminal via communication.

[0105] Step 7:

[0106] The terminal plays feedback messages received from the server using an audio output device and conveys them to the user. The input is the text data of the feedback. The output is audio data that the user can hear. This includes the specific operation of converting text feedback into speech using speech synthesis technology and playing it back in real time.

[0107] (Application Example 1)

[0108] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0109] Modern industrial settings demand efficient worker training and safety assurance. However, traditional training methods struggle to provide individualized training, particularly in promoting understanding of machine operation and safety procedures among new employees. Maintaining learning motivation is also difficult. This can lead to decreased work efficiency and increased safety risks.

[0110] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0111] In this invention, the server includes means for the user to request access to educational content via voice input, means for analyzing the voice input and dynamically selecting appropriate questions based on the user's history and skill level, and means for presenting quizzes related to the operation and safety procedures of industrial machinery and educating workers using an audio assistant. This enables workers to receive personalized education and efficiently acquire machine operation and safety procedures.

[0112] "A means by which a user requests access to educational content via voice input" refers to a method for a user to request educational content using their voice.

[0113] "A means of analyzing voice input and dynamically selecting appropriate problems based on the user's history and skill level" refers to a method of analyzing voice data to select tasks in real time that correspond to the user's past learning record and skill level.

[0114] "Means of presenting selected questions to the user via an audio output device" refers to a method of presenting selected quizzes or tasks to the user via audio.

[0115] "A means of receiving a user's verbal response and converting it into corresponding text data using speech recognition technology" refers to a technology that receives a user's voice response and converts it into text.

[0116] "A means of evaluating text data, determining correctness, and generating immediate feedback" refers to a method of evaluating converted text, determining whether it is correct or incorrect, and providing rapid feedback.

[0117] "Means of providing generated feedback to the user in audio format" refers to methods of communicating evaluation results to the user in audio format.

[0118] "Means of recording user progress and analyzing data to improve future plans" refers to technologies that store learner performance and evaluate data to optimize future educational plans.

[0119] "A method of educating workers by presenting quizzes related to the operation and safety procedures of industrial machinery and using an audio assistant" refers to a method of training workers by presenting questions related to the handling and safety measures of industrial equipment and utilizing an audio assistant.

[0120] The system for realizing this application primarily consists of three components working together: a server, terminals (including robots), and the user. The server resides in the cloud and is equipped with a speech recognition engine and a generative AI model. The server receives voice input from the user, converts it to text, and then dynamically selects appropriate quiz questions, taking into account the user's learning history and skill level.

[0121] The terminal includes industrial robots and voice output devices, enabling voice-based interaction. When a user requests access to educational content by voice, the terminal presents the user with a quiz received from the server. The user's voice response is sent back to the server by the terminal, where it is judged as correct or incorrect, and feedback is provided to the user by voice.

[0122] Users can use this system to advance their workplace training. For example, when learning how to operate industrial machinery, the robot might ask a quiz question like, "What is the first step in the safety procedure for this machine?" If the user answers, "Turn off the power," the server will determine if the answer is correct, and the robot will immediately provide feedback such as, "That's correct. Next, let's check the emergency stop button."

[0123] This coordinated operation of the entire system enables efficient worker training. An example of a prompt message might be: "We are considering how to use this quiz assistant to educate workers on emergency stop procedures for safe machine operation. Please describe specific question examples and the feedback process."

[0124] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0125] Step 1:

[0126] The device receives voice input from the user. The device uses its built-in microphone to capture the user's voice request as a digital audio signal. This input is processed as a request to access educational content.

[0127] Step 2:

[0128] The device sends voice data to the server. The voice data is sent to the server as data packets over the internet. Based on this input, the server uses a speech recognition engine to convert the data into text data.

[0129] Step 3:

[0130] The server performs speech recognition. Using the Google Cloud Speech-to-Text API, it converts the speech data into text and generates user requests in text format. This output allows the server to understand what the user is requesting.

[0131] Step 4:

[0132] The server selects appropriate questions based on the user's history and skill level. It analyzes text data, queries the user's learning history database, and dynamically selects the most suitable quiz questions. This data processing ensures that appropriate educational content is prepared for the user.

[0133] Step 5:

[0134] The server sends a selected question to the terminal. The selected question is sent to the terminal as a data packet and presented to the user via an audio output device. This output allows the user to take on the quiz.

[0135] Step 6:

[0136] The user answers the quiz using their voice. The user answers the quiz using their voice, and this is captured again as an audio signal by the device.

[0137] Step 7:

[0138] The device sends the user's voice response to the server. The voice data is then sent back to the server in digital format and used as input for speech recognition technology to convert it into text data.

[0139] Step 8:

[0140] The server analyzes the user's response and determines whether it is correct or incorrect. The converted text is compared to the correct answer data for the question, and an AI model determines whether it is correct or not. Feedback is generated as the output of this data calculation.

[0141] Step 9:

[0142] The server generates feedback and sends it to the terminal. The generated feedback is sent to the terminal as voice instructions, providing the user with the results verbally. This output allows the user to know the accuracy of their answers.

[0143] Step 10:

[0144] Users receive feedback and prepare for the next challenge. Based on the feedback, they confirm their understanding and prepare to continue learning.

[0145] Through these steps, the system can provide users with efficient education and immediate feedback.

[0146] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0147] The present invention is a learning support system that incorporates an emotion engine to enable users to learn more effectively. This system monitors the user's learning progress in real time through interaction with the server, terminal, and user. In the process, it recognizes the user's emotional state and appropriately adjusts the learning content and feedback.

[0148] The server runs a program that integrates speech recognition technology and an emotion engine. First, the server processes the user's voice data received from the terminal and converts the content of the voice into text. During this process, it uses the emotion engine to infer the user's emotions from the tone and intonation of the voice. This emotion information is used to select learning content and adjust feedback. The server also uses the emotion data to generate content that reduces user stress and to create customized messages that increase motivation.

[0149] The device acts as an interface to provide learners with optimal quiz questions and feedback based on information obtained through emotion recognition. For example, if the device determines that the user is experiencing high stress levels, it can offer relaxing music or guidance. Furthermore, if the user shows positive emotions toward their learning goals, the device may present more challenging questions to encourage further effort.

[0150] Users participate in quizzes through their devices and answer by voice. During this process, the user's speech is sent to a server for speech recognition and sentiment analysis. For example, if a user appears bored while answering easy questions, the system will detect this through sentiment analysis and attempt to stimulate their motivation by providing new, more challenging questions. Furthermore, if a user is facing a difficult question, the device will play encouraging messages to support them.

[0151] This system allows users to receive appropriate learning tailored to their emotional state, enabling them to learn more effectively and efficiently. Furthermore, by utilizing an emotion engine, a more personalized learning experience is provided.

[0152] The following describes the processing flow.

[0153] Step 1:

[0154] The user gives a voice command to the device saying, "Start the quiz."

[0155] Step 2:

[0156] The terminal converts the user's voice commands into digital signals and sends an initiation request to the server.

[0157] Step 3:

[0158] The server receives requests from the terminal and checks the user's learning history and profile.

[0159] Step 4:

[0160] The server activates speech recognition and emotion engines, analyzing the user's speech data to infer their emotional state. This allows it to understand the user's current emotions.

[0161] Step 5:

[0162] The server selects the most suitable quiz questions from the database based on the user's learning history and emotional state.

[0163] Step 6:

[0164] The server sends the selected quiz questions to the terminal, allowing it to prepare them for presentation to the user.

[0165] Step 7:

[0166] The device converts the received quiz question into audio and asks the user, "Here's the next question: What is the capital of Japan?"

[0167] Step 8:

[0168] The user answers the quiz verbally, responding with "It's Tokyo."

[0169] Step 9:

[0170] The terminal receives the user's voice response and sends it to the server as a digital signal.

[0171] Step 10:

[0172] The server uses speech recognition technology to convert speech into text and evaluates the response by comparing it to correct answer data.

[0173] Step 11:

[0174] The server then performs sentiment analysis on the text data to analyze changes in emotions during the response.

[0175] Step 12:

[0176] The server generates customized feedback for the user based on the results of correct / incorrect judgment and sentiment analysis. For example, if the answer is correct, it might say, "That's correct, congratulations!" and if it's incorrect, it might say, "That's unfortunate, but let's try again next time."

[0177] Step 13:

[0178] The server provides the generated feedback to the user as audio output via the terminal.

[0179] Step 14:

[0180] The terminal communicates feedback from the server to the user via voice, helping to enhance the user's motivation to learn.

[0181] Through this series of processes, the system can grasp the user's emotional state in real time and highly personalize the learning experience.

[0182] (Example 2)

[0183] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0184] The problem that this invention aims to solve is to enable learners to receive an appropriate learning experience that is tailored to their emotional state. Conventional learning support systems have not taken into account the user's emotions when adjusting content, making it difficult to maximize learning efficiency. Furthermore, the learning content is uniform, and there is little feedback or content provision tailored to the individual interests and emotional state of the learners. As a result, it has been difficult to effectively maintain the learners' concentration and motivation.

[0185] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0186] This invention includes a server that analyzes voice input and dynamically selects appropriate problems based on the learner's past learning information and ability level; a server that uses emotion analysis technology to infer the user's emotions from the tone and intonation of the voice and adjusts the selection of learning materials and feedback; and a server that generates personalized messages for stress reduction and motivation enhancement based on the user's emotion data. This makes it possible to provide personalized learning content that reflects the user's emotions in real time.

[0187] "Voice input" refers to the act of a user communicating information or instructions to a system via voice.

[0188] A "learner" is someone who uses a system for the purpose of acquiring knowledge or skills.

[0189] "Past learning information" refers to records of learning content and results that learners have worked on in the past.

[0190] "Competency level" is an indicator that shows the degree of a learner's knowledge or skills.

[0191] "Dynamic selection" refers to making the optimal choice in real time based on the situation and conditions.

[0192] "Emotion analysis technology" is a technology that uses voice and other data to infer a user's emotional state.

[0193] "Voice tone and intonation" refers to the qualitative characteristics of a voice, such as pitch and intonation.

[0194] "Learning materials" refers to all learning content provided to learners.

[0195] "Feedback" refers to evaluations and guidance provided based on a learner's performance and progress.

[0196] "Stress reduction" refers to support aimed at alleviating the burden and tension of learners.

[0197] "Motivation enhancement" refers to methods and initiatives aimed at increasing learners' motivation.

[0198] A "personalized message" is a special message tailored to each user's situation and needs.

[0199] "Real-time" refers to a situation where processing and responses are performed instantly in response to the current circumstances.

[0200] "Personalized learning content" refers to learning materials optimized based on the individual characteristics and needs of each learner.

[0201] This invention is an advanced system for learning support that analyzes the user's emotional state in real time and provides personalized learning content accordingly.

[0202] The server first receives voice input from the user via the terminal. This is done using a microphone or voice input-enabled device. The received voice data is converted into text data using the Google Cloud Speech-to-Text API. Next, Affectiva's sentiment analysis technology is used to infer the user's emotions from the tone and intonation of the voice. These processes are carried out using computing resources within the data center.

[0203] The device's role is to provide users with optimal quiz questions and feedback based on emotional information and learning content provided by the server. For example, it can present application problems using mathematical formulas as quiz questions via audio. This allows users to progress through learning effectively.

[0204] Users respond to presented quizzes and feedback verbally. This audio is then sent back to the server for speech recognition and sentiment analysis.

[0205] As a concrete example, the following prompt can be entered into a generative AI model: "If the AI ​​detects heightened emotions from the voice of a user participating in the quiz, what kind of challenging question should be presented as the next learning content?"

[0206] This system allows users to find the learning method best suited to their emotional state, enabling them to learn more focused and efficiently than with traditional methods.

[0207] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0208] Step 1:

[0209] The server receives voice input from the user. It uses the Google Cloud Speech-to-Text API to convert the voice data, which is sent via the terminal, into text data. This conversion results in the voice data being output as text. Specifically, if the user says, "This is a simple quiz," the audio is reproduced exactly as it is in text.

[0210] Step 2:

[0211] The server uses Affectiva's sentiment analysis technology to infer the user's emotions based on the converted text data. It performs sentiment analysis based on the tone and intonation of the input voice and outputs the user's emotional state. Specifically, this includes a process where a bright tone of voice input is interpreted as "excited."

[0212] Step 3:

[0213] Based on the results of sentiment analysis, the server selects appropriate learning content, taking into account the user's past learning history and ability level. Using sentiment data and learning history data as input, it outputs optimal quiz questions and learning materials. For example, if the user is excited, it may select more difficult questions.

[0214] Step 4:

[0215] The terminal provides learning content presented by the server to the user. It prepares to present quiz questions and other content to the user via audio using its voice output function. The input is the learning content, and the audio output to the user is the output of this step. A concrete example would be a scenario where applied math problems are read aloud.

[0216] Step 5:

[0217] The user responds to the presented question using voice. They then input their voice again into the device, and this data is sent to the server. This serves as input for the next voice recognition step.

[0218] Step 6:

[0219] The server converts the user's voice response back into text data and performs a correctness check. It takes the response, converted into text by speech recognition, as input, determines whether it's correct, and generates feedback. This process outputs voice feedback such as "That's correct" for correct answers.

[0220] Step 7:

[0221] The server generates feedback and, taking into account the user's emotional state, devises additional support messages. The input is the previously obtained sentiment analysis data and correct / incorrect judgment results, while the output is a personalized message of encouragement or advice. For example, it might include a message like, "Let's try a more fun problem next time!"

[0222] Step 8:

[0223] The device delivers generated feedback and additional messages to the user via voice. The output of this step is delivered directly to the user's ears through voice output. Through these steps, the user receives a learning experience optimized for them.

[0224] (Application Example 2)

[0225] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0226] There is a need to address the issue of not being able to provide learners with appropriate feedback and learning experiences that take their emotional state into account when they use digital learning materials. This invention aims to improve learning efficiency and motivation by analyzing the learner's emotions through voice input and providing optimal learning content and feedback according to that state.

[0227] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0228] In this invention, the server includes means for the user to request access to digital learning materials through voice input, means for analyzing the voice input and dynamically selecting appropriate problems based on the learner's emotional state and knowledge level, and means for generating personalized messages to enhance learning motivation based on the user's emotional analysis. This enables a personalized learning experience that reflects the learner's emotional state.

[0229] An "educational plan" refers to a set of procedures and policies that take into account the goals and progress that learners should achieve, and that enable them to learn efficiently.

[0230] "Digital learning materials" refer to educational content provided using electronic media, and include learning materials in various formats such as text, video, and audio.

[0231] "Emotional state" refers to the learner's psychological and emotional condition and is a factor that influences the effectiveness of learning.

[0232] A "personalized learning experience" refers to a learning experience tailored to individual needs by providing learning content and feedback optimized according to each learner's characteristics and emotional state.

[0233] "Feedback" refers to information provided to learners that shows results and reactions, and is used to promote learning improvement and continuous growth.

[0234] The system implementing this invention consists of a server incorporating a digital learning support program using an emotion engine, and a terminal used by the user. First, the user makes a voice input through the terminal and requests access to digital learning materials. This request is sent to the server, which uses speech recognition technology to convert the voice into text information and performs emotion analysis.

[0235] The server infers the user's emotional state from the converted text information and dynamically generates questions and feedback tailored to the individual learner's characteristics based on the results. Using the emotion engine, if a learner is experiencing stress, it can present content to reduce stress or generate personalized messages to boost motivation.

[0236] The device presents feedback and learning content received from the server to the user via an audio output device. The user's verbal response is then sent back to the server, and feedback is provided immediately. This system allows the user to receive appropriate learning tailored to their emotional state.

[0237] For example, if a learner shows signs of anxiety while tackling a complex problem, the server will generate and output a message such as, "Don't rush, try again." An example of a prompt message for the generative AI model could be, "If the user is feeling anxious during learning, please provide an encouraging message."

[0238] The hardware includes a microphone and audio output device, while the software includes a speech recognition library and sentiment analysis algorithms. This combination makes it possible to provide learners with a personalized learning experience.

[0239] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0240] Step 1:

[0241] The user provides voice input to the device. The input voice data is captured via the device's microphone and converted into a digital signal. This data is then sent to the server.

[0242] Step 2:

[0243] The server uses speech recognition technology to convert the received audio data into text data. Using a speech recognition library, it analyzes the content of the audio and generates text as linguistic information. This text data then serves as input for sentiment analysis.

[0244] Step 3:

[0245] The server uses an emotion engine to analyze the user's emotional state from text data. This process employs an emotion analysis algorithm that outputs an emotion label (e.g., joy, sadness, boredom) based on the tone and wording of the text. This emotional information is then used in the next step.

[0246] Step 4:

[0247] The server selects the optimal learning content based on emotional information and the user's knowledge level. Here, a generative AI model is used to dynamically generate personalized questions and feedback according to the emotional labels. This generated content is output and sent to the terminal.

[0248] Step 5:

[0249] The terminal presents feedback and learning content received from the server to the user via an audio output device. The user verbally answers the presented questions, and this becomes the next input.

[0250] Step 6:

[0251] The user's verbal response is again captured as audio data on the device and sent to the server. The server uses speech recognition technology to convert the response into text and generates data to determine whether it is correct or incorrect.

[0252] Step 7:

[0253] The server evaluates the accuracy of the answer and immediately generates feedback. The generated feedback is then output back to the terminal and provided to the user. This allows the user to receive learning progress and recommendations in real time.

[0254] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0255] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0256] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0257] [Second Embodiment]

[0258] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0259] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0260] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0261] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0262] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0263] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0264] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0265] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0266] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0267] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0268] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0269] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0270] This system provides a voice interface that enables users to learn using voice and offers a personalized learning experience. The following describes how the server, terminal, and user interact and how the system operates.

[0271] The server consists of a core database and an AI speech recognition engine. The server first references the user's learning history and skill level, and then selects appropriate quiz questions from the database. In this selection process, the server automatically chooses the most suitable questions based on the user's unique learning patterns and past performance. Furthermore, upon receiving the user's voice response, the server converts it into text data via the speech recognition engine and performs a correct / incorrect judgment. Based on the judgment result, it generates feedback for the user and sends it to their terminal.

[0272] The terminal includes a voice output device and a voice input device that directly interact with the user. It converts the user's voice input into a digital signal and transmits it to the server. It also provides the user with received quiz questions via voice and prompts them to answer. The terminal assists in converting voice to text using voice recognition technology, providing data for the server to perform accurate evaluations. Furthermore, it plays back feedback sent from the server via voice, conveying the user's learning progress.

[0273] Users participate in the quiz through interaction with their device. For example, if a user answers the question "What is the capital of Japan?" with "Tokyo," the server analyzes the audio and immediately provides feedback that the answer is correct. This method allows users to understand the accuracy of their learning in real time and move on to the next question. Users can also check their progress and past performance through the app on their device or the web interface, and adjust their learning direction accordingly.

[0274] In this way, this system utilizes speech recognition technology and digital communication technology to provide a highly personalized learning experience. Users can learn at their own pace, and effective learning is possible even while on the go or working.

[0275] The following describes the processing flow.

[0276] Step 1:

[0277] The user gives an instruction to start learning by voice to the terminal. For example, the user says "Start the quiz".

[0278] Step 2:

[0279] The terminal converts the user's voice instruction into a digital signal and sends it to the server as a request.

[0280] Step 3:

[0281] The server analyzes the received request and refers to the past performance and learning topic history from the user's learning history database.

[0282] Step 4:

[0283] Based on the learning history and the current skill level, the server selects appropriate quiz questions from the database.

[0284] Step 5:

[0285] The server sends the selected quiz questions to the terminal and instructs it to present them to the user.

[0286] Step 6:

[0287] The terminal vocalizes the received quiz questions using speech synthesis technology and asks the user.

[0288] Step 7:

[0289] The user answers the quiz questions by voice, and the terminal receives the speech.

[0290] Step 8:

[0291] The terminal sends the user's voice answer to the server as a digital signal. <00009​​​ The server uses speech recognition technology to convert the audio signal into text data and compares it to correct answer data to evaluate the response.

[0294] Step 10:

[0295] The server performs a correct / incorrect judgment, generates a feedback message, and sends it to the terminal.

[0296] Step 11:

[0297] The device provides feedback to the user via voice. For example, it might say, "That's correct!" or "That's incorrect, the correct answer is Tokyo."

[0298] Step 12:

[0299] The device records the user's answers and sends the data to the server to be used in selecting the next quiz.

[0300] In this way, the series of processes is completed, and the user can proceed to the next quiz.

[0301] (Example 1)

[0302] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0303] Modern learning systems often provide uniform content to all learners simultaneously, lacking individualization that takes into account each learner's skill level and learning pattern. This makes effective learning difficult and hinders the ability to maintain learner motivation.

[0304] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0305] In this invention, the server includes means for a user to request access to learning resources through voice input, means for analyzing the voice input and dynamically selecting appropriate questions based on the learner's history and skill level, and means for recording the user's learning progress and analyzing information to improve future plans. This provides an individualized learning experience tailored to the needs of each learner, enabling effective and motivating learning.

[0306] "Learning resources" refers to information or teaching materials used by users to advance their learning.

[0307] "Voice input" refers to the process by which a user provides information orally and the system receives it.

[0308] "Learner's history" refers to a record indicating what learning activities a user has carried out in the past.

[0309] "Skill level" refers to an indicator showing the degree of a user's current knowledge and ability.

[0310] "Dynamically select" refers to the process of selecting the most appropriate one in real-time or according to the situation.

[0311] "Voice output device" refers to a device that generates voice data from an electronic device in a form that can be heard by the user.

[0312] "Oral answer" refers to the act of a user providing an answer orally.

[0313] "Recognition technology" refers to technology for converting voice and other data into text or different formats.

[0314] "Corresponding text data" refers to text-formatted data converted from voice input.

[0315] "Correct / incorrect judgment" refers to the process of evaluating whether an input from a user is correct or incorrect.

[0316] "Feedback" refers to responses or guidance provided to users regarding their actions.

[0317] "Learning progress" refers to the state indicating how much progress a user has made through their learning process.

[0318] "Personalized experience" refers to a learning experience tailored to the individual needs and characteristics of each user.

[0319] In embodiments of the present invention, the system consists of a server, a terminal, and a user, each playing a specific role.

[0320] The server is the core information processing device of this invention. The server is equipped with a speech recognition engine and a database, and receives and processes user voice data internally. A common cloud-based speech recognition service can be used as the speech recognition engine; for example, the Google Cloud Speech-to-Text API can be used. The server first analyzes the voice input and selects the most suitable question by referencing the learner's history and skill level from the database. A generative AI model is used for this selection. The selected question is then sent from the server to the terminal.

[0321] A terminal is an information device that provides an interface with the user. The terminal is equipped with a voice input device and a voice output device, and transmits learning requests entered by the user via voice to the server. In addition, it plays back questions sent from the server via voice and presents them to the user. Furthermore, it records the user's answers and sends them back to the server. Specific forms of terminals include smartphones and dedicated tablets.

[0322] Users answer questions verbally via their devices, allowing them to learn while experiencing realistic interaction. For example, if a user receives the question "What is the capital of Japan?" and answers "Tokyo" verbally, the server analyzes the answer and immediately provides feedback on whether it is correct or incorrect.

[0323] As a concrete example, a prompt message for a generative AI model could be: "Based on the user's learning history and skill level, select the next question. Then, analyze the user's voice response and generate appropriate feedback."

[0324] In this way, by having the server, terminal, and user work together, it becomes possible to provide a real-time and personalized learning experience.

[0325] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0326] Step 1:

[0327] The server receives a voice access request from the user to learning content. It takes voice data from the device as input. To analyze this voice data, it uses a speech recognition engine to perform recognition processing and convert it into text data. This process involves receiving the voice signal as digital data and converting it into text format via a speech recognition API.

[0328] Step 2:

[0329] The server analyzes the converted text data and retrieves the user's learning history and skill level from the database. The input is audio data in text format. Based on this data, the server dynamically selects appropriate learning questions using a generative AI model. It performs database queries and considers past performance and learning patterns to determine the most suitable quiz.

[0330] Step 3:

[0331] The server sends the selected problem as text data to the terminal. The input is the data of the training problems selected by the AI ​​model. The output is the data that the terminal uses speech synthesis technology to provide to the user with the selected problem in voice. This involves sending the data via the HTTP protocol and the synthesis process to convert it into speech.

[0332] Step 4:

[0333] The user responds to questions presented by the terminal using voice. The input is the user's own voice response. The terminal records this voice, converts it into a digital signal, and sends it back to the server. This is where the microphone-based voice capture and digital conversion functions come into play.

[0334] Step 5:

[0335] The server converts the voice response data received from the terminal back into text data using a speech recognition engine. The input is digitized voice data. The server analyzes this data and determines whether it is correct or incorrect. The generated text data is then passed to an evaluation algorithm, which includes an analysis process to determine whether the quiz answers are correct or incorrect.

[0336] Step 6:

[0337] The server generates a feedback message based on the judgment result and sends it to the terminal. The input is the correct / incorrect judgment result. The output is the feedback content to be provided to the user in voice. The server then converts this feedback into text and sends it to the terminal via communication.

[0338] Step 7:

[0339] The terminal plays feedback messages received from the server using an audio output device and conveys them to the user. The input is the text data of the feedback. The output is audio data that the user can hear. This includes the specific operation of converting text feedback into speech using speech synthesis technology and playing it back in real time.

[0340] (Application Example 1)

[0341] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0342] Modern industrial settings demand efficient worker training and safety assurance. However, traditional training methods struggle to provide individualized training, particularly in promoting understanding of machine operation and safety procedures among new employees. Maintaining learning motivation is also difficult. This can lead to decreased work efficiency and increased safety risks.

[0343] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0344] In this invention, the server includes means for the user to request access to educational content via voice input, means for analyzing the voice input and dynamically selecting appropriate questions based on the user's history and skill level, and means for presenting quizzes related to the operation and safety procedures of industrial machinery and educating workers using an audio assistant. This enables workers to receive personalized education and efficiently acquire machine operation and safety procedures.

[0345] "A means by which a user requests access to educational content via voice input" refers to a method for a user to request educational content using their voice.

[0346] "A means of analyzing voice input and dynamically selecting appropriate problems based on the user's history and skill level" refers to a method of analyzing voice data to select tasks in real time that correspond to the user's past learning record and skill level.

[0347] "Means of presenting selected questions to the user via an audio output device" refers to a method of presenting selected quizzes or tasks to the user via audio.

[0348] "A means of receiving a user's verbal response and converting it into corresponding text data using speech recognition technology" refers to a technology that receives a user's voice response and converts it into text.

[0349] "A means of evaluating text data, determining correctness, and generating immediate feedback" refers to a method of evaluating converted text, determining whether it is correct or incorrect, and providing rapid feedback.

[0350] "Means of providing generated feedback to the user in audio format" refers to methods of communicating evaluation results to the user in audio format.

[0351] "Means of recording user progress and analyzing data to improve future plans" refers to technologies that store learner performance and evaluate data to optimize future educational plans.

[0352] "A method of educating workers by presenting quizzes related to the operation and safety procedures of industrial machinery and using an audio assistant" refers to a method of training workers by presenting questions related to the handling and safety measures of industrial equipment and utilizing an audio assistant.

[0353] The system for realizing this application primarily consists of three components working together: a server, terminals (including robots), and the user. The server resides in the cloud and is equipped with a speech recognition engine and a generative AI model. The server receives voice input from the user, converts it to text, and then dynamically selects appropriate quiz questions, taking into account the user's learning history and skill level.

[0354] The terminal includes industrial robots and voice output devices, enabling voice-based interaction. When a user requests access to educational content by voice, the terminal presents the user with a quiz received from the server. The user's voice response is sent back to the server by the terminal, where it is judged as correct or incorrect, and feedback is provided to the user by voice.

[0355] Users can use this system to advance their workplace training. For example, when learning how to operate industrial machinery, the robot might ask a quiz question like, "What is the first step in the safety procedure for this machine?" If the user answers, "Turn off the power," the server will determine if the answer is correct, and the robot will immediately provide feedback such as, "That's correct. Next, let's check the emergency stop button."

[0356] This coordinated operation of the entire system enables efficient worker training. An example of a prompt message might be: "We are considering how to use this quiz assistant to educate workers on emergency stop procedures for safe machine operation. Please describe specific question examples and the feedback process."

[0357] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0358] Step 1:

[0359] The device receives voice input from the user. The device uses its built-in microphone to capture the user's voice request as a digital audio signal. This input is processed as a request to access educational content.

[0360] Step 2:

[0361] The device sends voice data to the server. The voice data is sent to the server as data packets over the internet. Based on this input, the server uses a speech recognition engine to convert the data into text data.

[0362] Step 3:

[0363] The server performs speech recognition. Using the Google Cloud Speech-to-Text API, it converts the speech data into text and generates user requests in text format. This output allows the server to understand what the user is requesting.

[0364] Step 4:

[0365] The server selects appropriate questions based on the user's history and skill level. It analyzes text data, queries the user's learning history database, and dynamically selects the most suitable quiz questions. This data processing ensures that appropriate educational content is prepared for the user.

[0366] Step 5:

[0367] The server sends a selected question to the terminal. The selected question is sent to the terminal as a data packet and presented to the user via an audio output device. This output allows the user to take on the quiz.

[0368] Step 6:

[0369] The user answers the quiz using their voice. The user answers the quiz using their voice, and this is captured again as an audio signal by the device.

[0370] Step 7:

[0371] The device sends the user's voice response to the server. The voice data is then sent back to the server in digital format and used as input for speech recognition technology to convert it into text data.

[0372] Step 8:

[0373] The server analyzes the user's response and determines whether it is correct or incorrect. The converted text is compared to the correct answer data for the question, and an AI model determines whether it is correct or not. Feedback is generated as the output of this data calculation.

[0374] Step 9:

[0375] The server generates feedback and sends it to the terminal. The generated feedback is sent to the terminal as voice instructions, providing the user with the results verbally. This output allows the user to know the accuracy of their answers.

[0376] Step 10:

[0377] Users receive feedback and prepare for the next challenge. Based on the feedback, they confirm their understanding and prepare to continue learning.

[0378] Through these steps, the system can provide users with efficient education and immediate feedback.

[0379] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0380] The present invention is a learning support system that incorporates an emotion engine to enable users to learn more effectively. This system monitors the user's learning progress in real time through interaction with the server, terminal, and user. In the process, it recognizes the user's emotional state and appropriately adjusts the learning content and feedback.

[0381] The server runs a program that integrates speech recognition technology and an emotion engine. First, the server processes the user's voice data received from the terminal and converts the content of the voice into text. During this process, it uses the emotion engine to infer the user's emotions from the tone and intonation of the voice. This emotion information is used to select learning content and adjust feedback. The server also uses the emotion data to generate content that reduces user stress and to create customized messages that increase motivation.

[0382] The device acts as an interface to provide learners with optimal quiz questions and feedback based on information obtained through emotion recognition. For example, if the device determines that the user is experiencing high stress levels, it can offer relaxing music or guidance. Furthermore, if the user shows positive emotions toward their learning goals, the device may present more challenging questions to encourage further effort.

[0383] Users participate in quizzes through their devices and answer by voice. During this process, the user's speech is sent to a server for speech recognition and sentiment analysis. For example, if a user appears bored while answering easy questions, the system will detect this through sentiment analysis and attempt to stimulate their motivation by providing new, more challenging questions. Furthermore, if a user is facing a difficult question, the device will play encouraging messages to support them.

[0384] This system allows users to receive appropriate learning tailored to their emotional state, enabling them to learn more effectively and efficiently. Furthermore, by utilizing an emotion engine, a more personalized learning experience is provided.

[0385] The following describes the processing flow.

[0386] Step 1:

[0387] The user gives a voice command to the device saying, "Start the quiz."

[0388] Step 2:

[0389] The terminal converts the user's voice commands into digital signals and sends an initiation request to the server.

[0390] Step 3:

[0391] The server receives requests from the terminal and checks the user's learning history and profile.

[0392] Step 4:

[0393] The server activates speech recognition and emotion engines, analyzing the user's speech data to infer their emotional state. This allows it to understand the user's current emotions.

[0394] Step 5:

[0395] The server selects the most suitable quiz questions from the database based on the user's learning history and emotional state.

[0396] Step 6:

[0397] The server sends the selected quiz questions to the terminal, allowing it to prepare them for presentation to the user.

[0398] Step 7:

[0399] The device converts the received quiz question into audio and asks the user, "Here's the next question: What is the capital of Japan?"

[0400] Step 8:

[0401] The user answers the quiz verbally, responding with "It's Tokyo."

[0402] Step 9:

[0403] The terminal receives the user's voice response and sends it to the server as a digital signal.

[0404] Step 10:

[0405] The server uses speech recognition technology to convert speech into text and evaluates the response by comparing it to correct answer data.

[0406] Step 11:

[0407] The server then performs sentiment analysis on the text data to analyze changes in emotions during the response.

[0408] Step 12:

[0409] The server generates customized feedback for the user based on the results of correct / incorrect judgment and sentiment analysis. For example, if the answer is correct, it might say, "That's correct, congratulations!" and if it's incorrect, it might say, "That's unfortunate, but let's try again next time."

[0410] Step 13:

[0411] The server provides the generated feedback to the user as audio output via the terminal.

[0412] Step 14:

[0413] The terminal communicates feedback from the server to the user via voice, helping to enhance the user's motivation to learn.

[0414] Through this series of processes, the system can grasp the user's emotional state in real time and highly personalize the learning experience.

[0415] (Example 2)

[0416] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0417] The problem that this invention aims to solve is to enable learners to receive an appropriate learning experience that is tailored to their emotional state. Conventional learning support systems have not taken into account the user's emotions when adjusting content, making it difficult to maximize learning efficiency. Furthermore, the learning content is uniform, and there is little feedback or content provision tailored to the individual interests and emotional state of the learners. As a result, it has been difficult to effectively maintain the learners' concentration and motivation.

[0418] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0419] This invention includes a server that analyzes voice input and dynamically selects appropriate problems based on the learner's past learning information and ability level; a server that uses emotion analysis technology to infer the user's emotions from the tone and intonation of the voice and adjusts the selection of learning materials and feedback; and a server that generates personalized messages for stress reduction and motivation enhancement based on the user's emotion data. This makes it possible to provide personalized learning content that reflects the user's emotions in real time.

[0420] "Voice input" refers to the act of a user communicating information or instructions to a system via voice.

[0421] A "learner" is someone who uses a system for the purpose of acquiring knowledge or skills.

[0422] "Past learning information" refers to records of learning content and results that learners have worked on in the past.

[0423] "Competency level" is an indicator that shows the degree of a learner's knowledge or skills.

[0424] "Dynamic selection" refers to making the optimal choice in real time based on the situation and conditions.

[0425] "Emotion analysis technology" is a technology that uses voice and other data to infer a user's emotional state.

[0426] "Voice tone and intonation" refers to the qualitative characteristics of a voice, such as pitch and intonation.

[0427] "Learning materials" refers to all learning content provided to learners.

[0428] "Feedback" refers to evaluations and guidance provided based on a learner's performance and progress.

[0429] "Stress reduction" refers to support aimed at alleviating the burden and tension of learners.

[0430] "Motivation enhancement" refers to methods and initiatives aimed at increasing learners' motivation.

[0431] A "personalized message" is a special message tailored to each user's situation and needs.

[0432] "Real-time" refers to a situation where processing and responses are performed instantly in response to the current circumstances.

[0433] "Personalized learning content" refers to learning materials optimized based on the individual characteristics and needs of each learner.

[0434] This invention is an advanced system for learning support that analyzes the user's emotional state in real time and provides personalized learning content accordingly.

[0435] The server first receives voice input from the user via the terminal. This is done using a microphone or voice input-enabled device. The received voice data is converted into text data using the Google Cloud Speech-to-Text API. Next, Affectiva's sentiment analysis technology is used to infer the user's emotions from the tone and intonation of the voice. These processes are carried out using computing resources within the data center.

[0436] The device's role is to provide users with optimal quiz questions and feedback based on emotional information and learning content provided by the server. For example, it can present application problems using mathematical formulas as quiz questions via audio. This allows users to progress through learning effectively.

[0437] Users respond to presented quizzes and feedback verbally. This audio is then sent back to the server for speech recognition and sentiment analysis.

[0438] As a concrete example, the following prompt can be entered into a generative AI model: "If the AI ​​detects heightened emotions from the voice of a user participating in the quiz, what kind of challenging question should be presented as the next learning content?"

[0439] This system allows users to find the learning method best suited to their emotional state, enabling them to learn more focused and efficiently than with traditional methods.

[0440] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0441] Step 1:

[0442] The server receives voice input from the user. It uses the Google Cloud Speech-to-Text API to convert the voice data, which is sent via the terminal, into text data. This conversion results in the voice data being output as text. Specifically, if the user says, "This is a simple quiz," the audio is reproduced exactly as it is in text.

[0443] Step 2:

[0444] The server uses Affectiva's sentiment analysis technology to infer the user's emotions based on the converted text data. It performs sentiment analysis based on the tone and intonation of the input voice and outputs the user's emotional state. Specifically, this includes a process where a bright tone of voice input is interpreted as "excited."

[0445] Step 3:

[0446] Based on the results of sentiment analysis, the server selects appropriate learning content, taking into account the user's past learning history and ability level. Using sentiment data and learning history data as input, it outputs optimal quiz questions and learning materials. For example, if the user is excited, it may select more difficult questions.

[0447] Step 4:

[0448] The terminal provides learning content presented by the server to the user. It prepares to present quiz questions and other content to the user via audio using its voice output function. The input is the learning content, and the audio output to the user is the output of this step. A concrete example would be a scenario where applied math problems are read aloud.

[0449] Step 5:

[0450] The user responds to the presented question using voice. They then input their voice again into the device, and this data is sent to the server. This serves as input for the next voice recognition step.

[0451] Step 6:

[0452] The server converts the user's voice response back into text data and performs a correctness check. It takes the response, converted into text by speech recognition, as input, determines whether it's correct, and generates feedback. This process outputs voice feedback such as "That's correct" for correct answers.

[0453] Step 7:

[0454] The server generates feedback and, taking into account the user's emotional state, devises additional support messages. The input is the previously obtained sentiment analysis data and correct / incorrect judgment results, while the output is a personalized message of encouragement or advice. For example, it might include a message like, "Let's try a more fun problem next time!"

[0455] Step 8:

[0456] The device delivers generated feedback and additional messages to the user via voice. The output of this step is delivered directly to the user's ears through voice output. Through these steps, the user receives a learning experience optimized for them.

[0457] (Application Example 2)

[0458] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0459] There is a need to address the issue of not being able to provide learners with appropriate feedback and learning experiences that take their emotional state into account when they use digital learning materials. This invention aims to improve learning efficiency and motivation by analyzing the learner's emotions through voice input and providing optimal learning content and feedback according to that state.

[0460] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0461] In this invention, the server includes means for the user to request access to digital learning materials through voice input, means for analyzing the voice input and dynamically selecting appropriate problems based on the learner's emotional state and knowledge level, and means for generating personalized messages to enhance learning motivation based on the user's emotional analysis. This enables a personalized learning experience that reflects the learner's emotional state.

[0462] An "educational plan" refers to a set of procedures and policies that take into account the goals and progress that learners should achieve, and that enable them to learn efficiently.

[0463] "Digital learning materials" refer to educational content provided using electronic media, and include learning materials in various formats such as text, video, and audio.

[0464] "Emotional state" refers to the learner's psychological and emotional condition and is a factor that influences the effectiveness of learning.

[0465] A "personalized learning experience" refers to a learning experience tailored to individual needs by providing learning content and feedback optimized according to each learner's characteristics and emotional state.

[0466] "Feedback" refers to information provided to learners that shows results and reactions, and is used to promote learning improvement and continuous growth.

[0467] The system implementing this invention consists of a server incorporating a digital learning support program using an emotion engine, and a terminal used by the user. First, the user makes a voice input through the terminal and requests access to digital learning materials. This request is sent to the server, which uses speech recognition technology to convert the voice into text information and performs emotion analysis.

[0468] The server infers the user's emotional state from the converted text information and dynamically generates questions and feedback tailored to the individual learner's characteristics based on the results. Using the emotion engine, if a learner is experiencing stress, it can present content to reduce stress or generate personalized messages to boost motivation.

[0469] The device presents feedback and learning content received from the server to the user via an audio output device. The user's verbal response is then sent back to the server, and feedback is provided immediately. This system allows the user to receive appropriate learning tailored to their emotional state.

[0470] For example, if a learner shows signs of anxiety while tackling a complex problem, the server will generate and output a message such as, "Don't rush, try again." An example of a prompt message for the generative AI model could be, "If the user is feeling anxious during learning, please provide an encouraging message."

[0471] The hardware includes a microphone and audio output device, while the software includes a speech recognition library and sentiment analysis algorithms. This combination makes it possible to provide learners with a personalized learning experience.

[0472] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0473] Step 1:

[0474] The user provides voice input to the device. The input voice data is captured via the device's microphone and converted into a digital signal. This data is then sent to the server.

[0475] Step 2:

[0476] The server uses speech recognition technology to convert the received audio data into text data. Using a speech recognition library, it analyzes the content of the audio and generates text as linguistic information. This text data then serves as input for sentiment analysis.

[0477] Step 3:

[0478] The server uses an emotion engine to analyze the user's emotional state from text data. This process employs an emotion analysis algorithm that outputs an emotion label (e.g., joy, sadness, boredom) based on the tone and wording of the text. This emotional information is then used in the next step.

[0479] Step 4:

[0480] The server selects the optimal learning content based on emotional information and the user's knowledge level. Here, a generative AI model is used to dynamically generate personalized questions and feedback according to the emotional labels. This generated content is output and sent to the terminal.

[0481] Step 5:

[0482] The terminal presents feedback and learning content received from the server to the user via an audio output device. The user verbally answers the presented questions, and this becomes the next input.

[0483] Step 6:

[0484] The user's verbal response is again captured as audio data on the device and sent to the server. The server uses speech recognition technology to convert the response into text and generates data to determine whether it is correct or incorrect.

[0485] Step 7:

[0486] The server evaluates the accuracy of the answer and immediately generates feedback. The generated feedback is then output back to the terminal and provided to the user. This allows the user to receive learning progress and recommendations in real time.

[0487] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0488] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0489] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0490] [Third Embodiment]

[0491] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0492] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0493] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0494] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0495] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0496] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0497] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0498] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0499] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0500] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0501] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0502] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0503] This system provides a voice interface that enables users to learn using voice and offers a personalized learning experience. The following describes how the server, terminal, and user interact and how the system operates.

[0504] The server consists of a core database and an AI speech recognition engine. The server first references the user's learning history and skill level, and then selects appropriate quiz questions from the database. In this selection process, the server automatically chooses the most suitable questions based on the user's unique learning patterns and past performance. Furthermore, upon receiving the user's voice response, the server converts it into text data via the speech recognition engine and performs a correct / incorrect judgment. Based on the judgment result, it generates feedback for the user and sends it to their terminal.

[0505] The terminal includes a voice output device and a voice input device that directly interact with the user. It converts the user's voice input into a digital signal and transmits it to the server. It also provides the user with received quiz questions via voice and prompts them to answer. The terminal assists in converting voice to text using voice recognition technology, providing data for the server to perform accurate evaluations. Furthermore, it plays back feedback sent from the server via voice, conveying the user's learning progress.

[0506] Users participate in the quiz through interaction with their device. For example, if a user answers the question "What is the capital of Japan?" with "Tokyo," the server analyzes the audio and immediately provides feedback that the answer is correct. This method allows users to understand the accuracy of their learning in real time and move on to the next question. Users can also check their progress and past performance through the app on their device or the web interface, and adjust their learning direction accordingly.

[0507] In this way, this system utilizes speech recognition technology and digital communication technology to provide a highly personalized learning experience. Users can learn at their own pace, and effective learning is possible even while on the go or working.

[0508] The following describes the processing flow.

[0509] Step 1:

[0510] The user gives a voice command to the device to begin learning. For example, they might say, "Start the quiz."

[0511] Step 2:

[0512] The terminal converts the user's voice commands into digital signals and sends them to the server as requests.

[0513] Step 3:

[0514] The server parses the received request and references the user's learning history database for past performance and learning topics.

[0515] Step 4:

[0516] The server selects appropriate quiz questions from the database based on the learning history and current skill level.

[0517] Step 5:

[0518] The server sends the selected quiz questions to the terminal and instructs it to present them to the user.

[0519] Step 6:

[0520] The device uses speech synthesis technology to convert the received quiz questions into speech and presents them to the user.

[0521] Step 7:

[0522] The user answers quiz questions by voice, and the device receives their utterance.

[0523] Step 8:

[0524] The terminal transmits the user's voice response as a digital signal to the server.

[0525] Step 9:

[0526] The server uses speech recognition technology to convert the audio signal into text data and compares it to correct answer data to evaluate the response.

[0527] Step 10:

[0528] The server performs a correct / incorrect judgment, generates a feedback message, and sends it to the terminal.

[0529] Step 11:

[0530] The device provides feedback to the user via voice. For example, it might say, "That's correct!" or "That's incorrect, the correct answer is Tokyo."

[0531] Step 12:

[0532] The device records the user's answers and sends the data to the server to be used in selecting the next quiz.

[0533] In this way, the series of processes is completed, and the user can proceed to the next quiz.

[0534] (Example 1)

[0535] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0536] Modern learning systems often provide uniform content to all learners simultaneously, lacking individualization that takes into account each learner's skill level and learning pattern. This makes effective learning difficult and hinders the ability to maintain learner motivation.

[0537] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0538] In this invention, the server includes means for the user to request access to learning resources through voice input, means for analyzing the voice input and dynamically selecting appropriate problems based on the learner's history and skill level, and means for recording the user's learning progress and analyzing the information to improve future plans. This provides a personalized learning experience tailored to the individual needs of each learner, enabling effective and highly motivating learning.

[0539] "Learning resources" refer to information or learning materials that users use to advance their learning.

[0540] "Voice input" refers to the process in which a user provides information verbally and a system receives it.

[0541] "Learner history" refers to a record showing what learning activities a user has engaged in in the past.

[0542] "Skill level" refers to an indicator that shows the current level of a user's knowledge and abilities.

[0543] "Dynamic selection" refers to the process of choosing the most appropriate option in real time or based on the situation.

[0544] An "audio output device" refers to a device that generates audio data from an electronic device in a format that the user can listen to.

[0545] "Verbal response" refers to the act of a user providing an answer via voice.

[0546] "Recognition technology" refers to technologies for converting audio and other data into text or other formats.

[0547] "The relevant text data" refers to text data converted from voice input.

[0548] "Correct / incorrect judgment" refers to the process of evaluating whether user input is correct or incorrect.

[0549] "Feedback" refers to responses or guidance provided to users regarding their actions.

[0550] "Learning progress" refers to the state indicating how much progress a user has made through their learning process.

[0551] "Personalized experience" refers to a learning experience tailored to the individual needs and characteristics of each user.

[0552] In embodiments of the present invention, the system consists of a server, a terminal, and a user, each playing a specific role.

[0553] The server is the core information processing device of this invention. The server is equipped with a speech recognition engine and a database, and receives and processes user voice data internally. A common cloud-based speech recognition service can be used as the speech recognition engine; for example, the Google Cloud Speech-to-Text API can be used. The server first analyzes the voice input and selects the most suitable question by referencing the learner's history and skill level from the database. A generative AI model is used for this selection. The selected question is then sent from the server to the terminal.

[0554] A terminal is an information device that provides an interface with the user. The terminal is equipped with a voice input device and a voice output device, and transmits learning requests entered by the user via voice to the server. In addition, it plays back questions sent from the server via voice and presents them to the user. Furthermore, it records the user's answers and sends them back to the server. Specific forms of terminals include smartphones and dedicated tablets.

[0555] Users answer questions verbally via their devices, allowing them to learn while experiencing realistic interaction. For example, if a user receives the question "What is the capital of Japan?" and answers "Tokyo" verbally, the server analyzes the answer and immediately provides feedback on whether it is correct or incorrect.

[0556] As a concrete example, a prompt message for a generative AI model could be: "Based on the user's learning history and skill level, select the next question. Then, analyze the user's voice response and generate appropriate feedback."

[0557] In this way, by having the server, terminal, and user work together, it becomes possible to provide a real-time and personalized learning experience.

[0558] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0559] Step 1:

[0560] The server receives a voice access request from the user to learning content. It takes voice data from the device as input. To analyze this voice data, it uses a speech recognition engine to perform recognition processing and convert it into text data. This process involves receiving the voice signal as digital data and converting it into text format via a speech recognition API.

[0561] Step 2:

[0562] The server analyzes the converted text data and retrieves the user's learning history and skill level from the database. The input is audio data in text format. Based on this data, the server dynamically selects appropriate learning questions using a generative AI model. It performs database queries and considers past performance and learning patterns to determine the most suitable quiz.

[0563] Step 3:

[0564] The server sends the selected problem as text data to the terminal. The input is the data of the training problems selected by the AI ​​model. The output is the data that the terminal uses speech synthesis technology to provide to the user with the selected problem in voice. This involves sending the data via the HTTP protocol and the synthesis process to convert it into speech.

[0565] Step 4:

[0566] The user responds to questions presented by the terminal using voice. The input is the user's own voice response. The terminal records this voice, converts it into a digital signal, and sends it back to the server. This is where the microphone-based voice capture and digital conversion functions come into play.

[0567] Step 5:

[0568] The server converts the voice response data received from the terminal back into text data using a speech recognition engine. The input is digitized voice data. The server analyzes this data and determines whether it is correct or incorrect. The generated text data is then passed to an evaluation algorithm, which includes an analysis process to determine whether the quiz answers are correct or incorrect.

[0569] Step 6:

[0570] The server generates a feedback message based on the judgment result and sends it to the terminal. The input is the correct / incorrect judgment result. The output is the feedback content to be provided to the user in voice. The server then converts this feedback into text and sends it to the terminal via communication.

[0571] Step 7:

[0572] The terminal plays feedback messages received from the server using an audio output device and conveys them to the user. The input is the text data of the feedback. The output is audio data that the user can hear. This includes the specific operation of converting text feedback into speech using speech synthesis technology and playing it back in real time.

[0573] (Application Example 1)

[0574] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0575] Modern industrial settings demand efficient worker training and safety assurance. However, traditional training methods struggle to provide individualized training, particularly in promoting understanding of machine operation and safety procedures among new employees. Maintaining learning motivation is also difficult. This can lead to decreased work efficiency and increased safety risks.

[0576] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0577] In this invention, the server includes means for the user to request access to educational content via voice input, means for analyzing the voice input and dynamically selecting appropriate questions based on the user's history and skill level, and means for presenting quizzes related to the operation and safety procedures of industrial machinery and educating workers using an audio assistant. This enables workers to receive personalized education and efficiently acquire machine operation and safety procedures.

[0578] "A means by which a user requests access to educational content via voice input" refers to a method for a user to request educational content using their voice.

[0579] "A means of analyzing voice input and dynamically selecting appropriate problems based on the user's history and skill level" refers to a method of analyzing voice data to select tasks in real time that correspond to the user's past learning record and skill level.

[0580] "Means of presenting selected questions to the user via an audio output device" refers to a method of presenting selected quizzes or tasks to the user via audio.

[0581] "A means of receiving a user's verbal response and converting it into corresponding text data using speech recognition technology" refers to a technology that receives a user's voice response and converts it into text.

[0582] "A means of evaluating text data, determining correctness, and generating immediate feedback" refers to a method of evaluating converted text, determining whether it is correct or incorrect, and providing rapid feedback.

[0583] "Means of providing generated feedback to the user in audio format" refers to methods of communicating evaluation results to the user in audio format.

[0584] "Means of recording user progress and analyzing data to improve future plans" refers to technologies that store learner performance and evaluate data to optimize future educational plans.

[0585] "A method of educating workers by presenting quizzes related to the operation and safety procedures of industrial machinery and using an audio assistant" refers to a method of training workers by presenting questions related to the handling and safety measures of industrial equipment and utilizing an audio assistant.

[0586] The system for realizing this application primarily consists of three components working together: a server, terminals (including robots), and the user. The server resides in the cloud and is equipped with a speech recognition engine and a generative AI model. The server receives voice input from the user, converts it to text, and then dynamically selects appropriate quiz questions, taking into account the user's learning history and skill level.

[0587] The terminal includes industrial robots and voice output devices, enabling voice-based interaction. When a user requests access to educational content by voice, the terminal presents the user with a quiz received from the server. The user's voice response is sent back to the server by the terminal, where it is judged as correct or incorrect, and feedback is provided to the user by voice.

[0588] Users can use this system to advance their workplace training. For example, when learning how to operate industrial machinery, the robot might ask a quiz question like, "What is the first step in the safety procedure for this machine?" If the user answers, "Turn off the power," the server will determine if the answer is correct, and the robot will immediately provide feedback such as, "That's correct. Next, let's check the emergency stop button."

[0589] This coordinated operation of the entire system enables efficient worker training. An example of a prompt message might be: "We are considering how to use this quiz assistant to educate workers on emergency stop procedures for safe machine operation. Please describe specific question examples and the feedback process."

[0590] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0591] Step 1:

[0592] The device receives voice input from the user. The device uses its built-in microphone to capture the user's voice request as a digital audio signal. This input is processed as a request to access educational content.

[0593] Step 2:

[0594] The device sends voice data to the server. The voice data is sent to the server as data packets over the internet. Based on this input, the server uses a speech recognition engine to convert the data into text data.

[0595] Step 3:

[0596] The server performs speech recognition. Using the Google Cloud Speech-to-Text API, it converts the speech data into text and generates user requests in text format. This output allows the server to understand what the user is requesting.

[0597] Step 4:

[0598] The server selects appropriate questions based on the user's history and skill level. It analyzes text data, queries the user's learning history database, and dynamically selects the most suitable quiz questions. This data processing ensures that appropriate educational content is prepared for the user.

[0599] Step 5:

[0600] The server sends a selected question to the terminal. The selected question is sent to the terminal as a data packet and presented to the user via an audio output device. This output allows the user to take on the quiz.

[0601] Step 6:

[0602] The user answers the quiz using their voice. The user answers the quiz using their voice, and this is captured again as an audio signal by the device.

[0603] Step 7:

[0604] The device sends the user's voice response to the server. The voice data is then sent back to the server in digital format and used as input for speech recognition technology to convert it into text data.

[0605] Step 8:

[0606] The server analyzes the user's response and determines whether it is correct or incorrect. The converted text is compared to the correct answer data for the question, and an AI model determines whether it is correct or not. Feedback is generated as the output of this data calculation.

[0607] Step 9:

[0608] The server generates feedback and sends it to the terminal. The generated feedback is sent to the terminal as voice instructions, providing the user with the results verbally. This output allows the user to know the accuracy of their answers.

[0609] Step 10:

[0610] Users receive feedback and prepare for the next challenge. Based on the feedback, they confirm their understanding and prepare to continue learning.

[0611] Through these steps, the system can provide users with efficient education and immediate feedback.

[0612] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0613] The present invention is a learning support system that incorporates an emotion engine to enable users to learn more effectively. This system monitors the user's learning progress in real time through interaction with the server, terminal, and user. In the process, it recognizes the user's emotional state and appropriately adjusts the learning content and feedback.

[0614] The server runs a program that integrates speech recognition technology and an emotion engine. First, the server processes the user's voice data received from the terminal and converts the content of the voice into text. During this process, it uses the emotion engine to infer the user's emotions from the tone and intonation of the voice. This emotion information is used to select learning content and adjust feedback. The server also uses the emotion data to generate content that reduces user stress and to create customized messages that increase motivation.

[0615] The device acts as an interface to provide learners with optimal quiz questions and feedback based on information obtained through emotion recognition. For example, if the device determines that the user is experiencing high stress levels, it can offer relaxing music or guidance. Furthermore, if the user shows positive emotions toward their learning goals, the device may present more challenging questions to encourage further effort.

[0616] Users participate in quizzes through their devices and answer by voice. During this process, the user's speech is sent to a server for speech recognition and sentiment analysis. For example, if a user appears bored while answering easy questions, the system will detect this through sentiment analysis and attempt to stimulate their motivation by providing new, more challenging questions. Furthermore, if a user is facing a difficult question, the device will play encouraging messages to support them.

[0617] This system allows users to receive appropriate learning tailored to their emotional state, enabling them to learn more effectively and efficiently. Furthermore, by utilizing an emotion engine, a more personalized learning experience is provided.

[0618] The following describes the processing flow.

[0619] Step 1:

[0620] The user gives a voice command to the device saying, "Start the quiz."

[0621] Step 2:

[0622] The terminal converts the user's voice commands into digital signals and sends an initiation request to the server.

[0623] Step 3:

[0624] The server receives requests from the terminal and checks the user's learning history and profile.

[0625] Step 4:

[0626] The server activates speech recognition and emotion engines, analyzing the user's speech data to infer their emotional state. This allows it to understand the user's current emotions.

[0627] Step 5:

[0628] The server selects the most suitable quiz questions from the database based on the user's learning history and emotional state.

[0629] Step 6:

[0630] The server sends the selected quiz questions to the terminal, allowing it to prepare them for presentation to the user.

[0631] Step 7:

[0632] The device converts the received quiz question into audio and asks the user, "Here's the next question: What is the capital of Japan?"

[0633] Step 8:

[0634] The user answers the quiz verbally, responding with "It's Tokyo."

[0635] Step 9:

[0636] The terminal receives the user's voice response and sends it to the server as a digital signal.

[0637] Step 10:

[0638] The server uses speech recognition technology to convert speech into text and evaluates the response by comparing it to correct answer data.

[0639] Step 11:

[0640] The server then performs sentiment analysis on the text data to analyze changes in emotions during the response.

[0641] Step 12:

[0642] The server generates customized feedback for the user based on the results of correct / incorrect judgment and sentiment analysis. For example, if the answer is correct, it might say, "That's correct, congratulations!" and if it's incorrect, it might say, "That's unfortunate, but let's try again next time."

[0643] Step 13:

[0644] The server provides the generated feedback to the user as audio output via the terminal.

[0645] Step 14:

[0646] The terminal communicates feedback from the server to the user via voice, helping to enhance the user's motivation to learn.

[0647] Through this series of processes, the system can grasp the user's emotional state in real time and highly personalize the learning experience.

[0648] (Example 2)

[0649] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0650] The problem that this invention aims to solve is to enable learners to receive an appropriate learning experience that is tailored to their emotional state. Conventional learning support systems have not taken into account the user's emotions when adjusting content, making it difficult to maximize learning efficiency. Furthermore, the learning content is uniform, and there is little feedback or content provision tailored to the individual interests and emotional state of the learners. As a result, it has been difficult to effectively maintain the learners' concentration and motivation.

[0651] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0652] This invention includes a server that analyzes voice input and dynamically selects appropriate problems based on the learner's past learning information and ability level; a server that uses emotion analysis technology to infer the user's emotions from the tone and intonation of the voice and adjusts the selection of learning materials and feedback; and a server that generates personalized messages for stress reduction and motivation enhancement based on the user's emotion data. This makes it possible to provide personalized learning content that reflects the user's emotions in real time.

[0653] "Voice input" refers to the act of a user communicating information or instructions to a system via voice.

[0654] A "learner" is someone who uses a system for the purpose of acquiring knowledge or skills.

[0655] "Past learning information" refers to records of learning content and results that learners have worked on in the past.

[0656] "Competency level" is an indicator that shows the degree of a learner's knowledge or skills.

[0657] "Dynamic selection" refers to making the optimal choice in real time based on the situation and conditions.

[0658] "Emotion analysis technology" is a technology that uses voice and other data to infer a user's emotional state.

[0659] "Voice tone and intonation" refers to the qualitative characteristics of a voice, such as pitch and intonation.

[0660] "Learning materials" refers to all learning content provided to learners.

[0661] "Feedback" refers to evaluations and guidance provided based on a learner's performance and progress.

[0662] "Stress reduction" refers to support aimed at alleviating the burden and tension of learners.

[0663] "Motivation enhancement" refers to methods and initiatives aimed at increasing learners' motivation.

[0664] A "personalized message" is a special message tailored to each user's situation and needs.

[0665] "Real-time" refers to a situation where processing and responses are performed instantly in response to the current circumstances.

[0666] "Personalized learning content" refers to learning materials optimized based on the individual characteristics and needs of each learner.

[0667] This invention is an advanced system for learning support that analyzes the user's emotional state in real time and provides personalized learning content accordingly.

[0668] The server first receives voice input from the user via the terminal. This is done using a microphone or voice input-enabled device. The received voice data is converted into text data using the Google Cloud Speech-to-Text API. Next, Affectiva's sentiment analysis technology is used to infer the user's emotions from the tone and intonation of the voice. These processes are carried out using computing resources within the data center.

[0669] The device's role is to provide users with optimal quiz questions and feedback based on emotional information and learning content provided by the server. For example, it can present application problems using mathematical formulas as quiz questions via audio. This allows users to progress through learning effectively.

[0670] Users respond to presented quizzes and feedback verbally. This audio is then sent back to the server for speech recognition and sentiment analysis.

[0671] As a concrete example, the following prompt can be entered into a generative AI model: "If the AI ​​detects heightened emotions from the voice of a user participating in the quiz, what kind of challenging question should be presented as the next learning content?"

[0672] This system allows users to find the learning method best suited to their emotional state, enabling them to learn more focused and efficiently than with traditional methods.

[0673] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0674] Step 1:

[0675] The server receives voice input from the user. It uses the Google Cloud Speech-to-Text API to convert the voice data, which is sent via the terminal, into text data. This conversion results in the voice data being output as text. Specifically, if the user says, "This is a simple quiz," the audio is reproduced exactly as it is in text.

[0676] Step 2:

[0677] The server uses Affectiva's sentiment analysis technology to infer the user's emotions based on the converted text data. It performs sentiment analysis based on the tone and intonation of the input voice and outputs the user's emotional state. Specifically, this includes a process where a bright tone of voice input is interpreted as "excited."

[0678] Step 3:

[0679] Based on the results of sentiment analysis, the server selects appropriate learning content, taking into account the user's past learning history and ability level. Using sentiment data and learning history data as input, it outputs optimal quiz questions and learning materials. For example, if the user is excited, it may select more difficult questions.

[0680] Step 4:

[0681] The terminal provides learning content presented by the server to the user. It prepares to present quiz questions and other content to the user via audio using its voice output function. The input is the learning content, and the audio output to the user is the output of this step. A concrete example would be a scenario where applied math problems are read aloud.

[0682] Step 5:

[0683] The user responds to the presented question using voice. They then input their voice again into the device, and this data is sent to the server. This serves as input for the next voice recognition step.

[0684] Step 6:

[0685] The server converts the user's voice response back into text data and performs a correctness check. It takes the response, converted into text by speech recognition, as input, determines whether it's correct, and generates feedback. This process outputs voice feedback such as "That's correct" for correct answers.

[0686] Step 7:

[0687] The server generates feedback and, taking into account the user's emotional state, devises additional support messages. The input is the previously obtained sentiment analysis data and correct / incorrect judgment results, while the output is a personalized message of encouragement or advice. For example, it might include a message like, "Let's try a more fun problem next time!"

[0688] Step 8:

[0689] The device delivers generated feedback and additional messages to the user via voice. The output of this step is delivered directly to the user's ears through voice output. Through these steps, the user receives a learning experience optimized for them.

[0690] (Application Example 2)

[0691] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0692] There is a need to address the issue of not being able to provide learners with appropriate feedback and learning experiences that take their emotional state into account when they use digital learning materials. This invention aims to improve learning efficiency and motivation by analyzing the learner's emotions through voice input and providing optimal learning content and feedback according to that state.

[0693] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0694] In this invention, the server includes means for the user to request access to digital learning materials through voice input, means for analyzing the voice input and dynamically selecting appropriate problems based on the learner's emotional state and knowledge level, and means for generating personalized messages to enhance learning motivation based on the user's emotional analysis. This enables a personalized learning experience that reflects the learner's emotional state.

[0695] An "educational plan" refers to a set of procedures and policies that take into account the goals and progress that learners should achieve, and that enable them to learn efficiently.

[0696] "Digital learning materials" refer to educational content provided using electronic media, and include learning materials in various formats such as text, video, and audio.

[0697] "Emotional state" refers to the learner's psychological and emotional condition and is a factor that influences the effectiveness of learning.

[0698] A "personalized learning experience" refers to a learning experience tailored to individual needs by providing learning content and feedback optimized according to each learner's characteristics and emotional state.

[0699] "Feedback" refers to information provided to learners that shows results and reactions, and is used to promote learning improvement and continuous growth.

[0700] The system implementing this invention consists of a server incorporating a digital learning support program using an emotion engine, and a terminal used by the user. First, the user makes a voice input through the terminal and requests access to digital learning materials. This request is sent to the server, which uses speech recognition technology to convert the voice into text information and performs emotion analysis.

[0701] The server infers the user's emotional state from the converted text information and dynamically generates questions and feedback tailored to the individual learner's characteristics based on the results. Using the emotion engine, if a learner is experiencing stress, it can present content to reduce stress or generate personalized messages to boost motivation.

[0702] The device presents feedback and learning content received from the server to the user via an audio output device. The user's verbal response is then sent back to the server, and feedback is provided immediately. This system allows the user to receive appropriate learning tailored to their emotional state.

[0703] For example, if a learner shows signs of anxiety while tackling a complex problem, the server will generate and output a message such as, "Don't rush, try again." An example of a prompt message for the generative AI model could be, "If the user is feeling anxious during learning, please provide an encouraging message."

[0704] The hardware includes a microphone and audio output device, while the software includes a speech recognition library and sentiment analysis algorithms. This combination makes it possible to provide learners with a personalized learning experience.

[0705] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0706] Step 1:

[0707] The user provides voice input to the device. The input voice data is captured via the device's microphone and converted into a digital signal. This data is then sent to the server.

[0708] Step 2:

[0709] The server uses speech recognition technology to convert the received audio data into text data. Using a speech recognition library, it analyzes the content of the audio and generates text as linguistic information. This text data then serves as input for sentiment analysis.

[0710] Step 3:

[0711] The server uses an emotion engine to analyze the user's emotional state from text data. This process employs an emotion analysis algorithm that outputs an emotion label (e.g., joy, sadness, boredom) based on the tone and wording of the text. This emotional information is then used in the next step.

[0712] Step 4:

[0713] The server selects the optimal learning content based on emotional information and the user's knowledge level. Here, a generative AI model is used to dynamically generate personalized questions and feedback according to the emotional labels. This generated content is output and sent to the terminal.

[0714] Step 5:

[0715] The terminal presents feedback and learning content received from the server to the user via an audio output device. The user verbally answers the presented questions, and this becomes the next input.

[0716] Step 6:

[0717] The user's verbal response is again captured as audio data on the device and sent to the server. The server uses speech recognition technology to convert the response into text and generates data to determine whether it is correct or incorrect.

[0718] Step 7:

[0719] The server evaluates the accuracy of the answer and immediately generates feedback. The generated feedback is then output back to the terminal and provided to the user. This allows the user to receive learning progress and recommendations in real time.

[0720] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0721] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0722] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0723] [Fourth Embodiment]

[0724] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0725] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0726] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0727] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0728] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0729] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0730] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0731] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0732] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0733] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0734] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0735] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0736] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0737] This system provides a voice interface that enables users to learn using voice and offers a personalized learning experience. The following describes how the server, terminal, and user interact and how the system operates.

[0738] The server consists of a core database and an AI speech recognition engine. The server first references the user's learning history and skill level, and then selects appropriate quiz questions from the database. In this selection process, the server automatically chooses the most suitable questions based on the user's unique learning patterns and past performance. Furthermore, upon receiving the user's voice response, the server converts it into text data via the speech recognition engine and performs a correct / incorrect judgment. Based on the judgment result, it generates feedback for the user and sends it to their terminal.

[0739] The terminal includes a voice output device and a voice input device that directly interact with the user. It converts the user's voice input into a digital signal and transmits it to the server. It also provides the user with received quiz questions via voice and prompts them to answer. The terminal assists in converting voice to text using voice recognition technology, providing data for the server to perform accurate evaluations. Furthermore, it plays back feedback sent from the server via voice, conveying the user's learning progress.

[0740] Users participate in the quiz through interaction with their device. For example, if a user answers the question "What is the capital of Japan?" with "Tokyo," the server analyzes the audio and immediately provides feedback that the answer is correct. This method allows users to understand the accuracy of their learning in real time and move on to the next question. Users can also check their progress and past performance through the app on their device or the web interface, and adjust their learning direction accordingly.

[0741] In this way, this system utilizes speech recognition technology and digital communication technology to provide a highly personalized learning experience. Users can learn at their own pace, and effective learning is possible even while on the go or working.

[0742] The following describes the processing flow.

[0743] Step 1:

[0744] The user gives a voice command to the device to begin learning. For example, they might say, "Start the quiz."

[0745] Step 2:

[0746] The terminal converts the user's voice commands into digital signals and sends them to the server as requests.

[0747] Step 3:

[0748] The server parses the received request and references the user's learning history database for past performance and learning topics.

[0749] Step 4:

[0750] The server selects appropriate quiz questions from the database based on the learning history and current skill level.

[0751] Step 5:

[0752] The server sends the selected quiz questions to the terminal and instructs it to present them to the user.

[0753] Step 6:

[0754] The device uses speech synthesis technology to convert the received quiz questions into speech and presents them to the user.

[0755] Step 7:

[0756] The user answers quiz questions by voice, and the device receives their utterance.

[0757] Step 8:

[0758] The terminal transmits the user's voice response as a digital signal to the server.

[0759] Step 9:

[0760] The server uses speech recognition technology to convert the audio signal into text data and compares it to correct answer data to evaluate the response.

[0761] Step 10:

[0762] The server performs a correct / incorrect judgment, generates a feedback message, and sends it to the terminal.

[0763] Step 11:

[0764] The device provides feedback to the user via voice. For example, it might say, "That's correct!" or "That's incorrect, the correct answer is Tokyo."

[0765] Step 12:

[0766] The device records the user's answers and sends the data to the server to be used in selecting the next quiz.

[0767] In this way, the series of processes is completed, and the user can proceed to the next quiz.

[0768] (Example 1)

[0769] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0770] Modern learning systems often provide uniform content to all learners simultaneously, lacking individualization that takes into account each learner's skill level and learning pattern. This makes effective learning difficult and hinders the ability to maintain learner motivation.

[0771] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0772] In this invention, the server includes means for the user to request access to learning resources through voice input, means for analyzing the voice input and dynamically selecting appropriate problems based on the learner's history and skill level, and means for recording the user's learning progress and analyzing the information to improve future plans. This provides a personalized learning experience tailored to the individual needs of each learner, enabling effective and highly motivating learning.

[0773] "Learning resources" refer to information or learning materials that users use to advance their learning.

[0774] "Voice input" refers to the process in which a user provides information verbally and a system receives it.

[0775] "Learner history" refers to a record showing what learning activities a user has engaged in in the past.

[0776] "Skill level" refers to an indicator that shows the current level of a user's knowledge and abilities.

[0777] "Dynamic selection" refers to the process of choosing the most appropriate option in real time or based on the situation.

[0778] An "audio output device" refers to a device that generates audio data from an electronic device in a format that the user can listen to.

[0779] "Verbal response" refers to the act of a user providing an answer via voice.

[0780] "Recognition technology" refers to technologies for converting audio and other data into text or other formats.

[0781] "The relevant text data" refers to text data converted from voice input.

[0782] "Correct / incorrect judgment" refers to the process of evaluating whether user input is correct or incorrect.

[0783] "Feedback" refers to responses or guidance provided to users regarding their actions.

[0784] "Learning progress" refers to the state indicating how much progress a user has made through their learning process.

[0785] "Personalized experience" refers to a learning experience tailored to the individual needs and characteristics of each user.

[0786] In embodiments of the present invention, the system consists of a server, a terminal, and a user, each playing a specific role.

[0787] The server is the core information processing device of this invention. The server is equipped with a speech recognition engine and a database, and receives and processes user voice data internally. A common cloud-based speech recognition service can be used as the speech recognition engine; for example, the Google Cloud Speech-to-Text API can be used. The server first analyzes the voice input and selects the most suitable question by referencing the learner's history and skill level from the database. A generative AI model is used for this selection. The selected question is then sent from the server to the terminal.

[0788] A terminal is an information device that provides an interface with the user. The terminal is equipped with a voice input device and a voice output device, and transmits learning requests entered by the user via voice to the server. In addition, it plays back questions sent from the server via voice and presents them to the user. Furthermore, it records the user's answers and sends them back to the server. Specific forms of terminals include smartphones and dedicated tablets.

[0789] Users answer questions verbally via their devices, allowing them to learn while experiencing realistic interaction. For example, if a user receives the question "What is the capital of Japan?" and answers "Tokyo" verbally, the server analyzes the answer and immediately provides feedback on whether it is correct or incorrect.

[0790] As a concrete example, a prompt message for a generative AI model could be: "Based on the user's learning history and skill level, select the next question. Then, analyze the user's voice response and generate appropriate feedback."

[0791] In this way, by having the server, terminal, and user work together, it becomes possible to provide a real-time and personalized learning experience.

[0792] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0793] Step 1:

[0794] The server receives a voice access request from the user to learning content. It takes voice data from the device as input. To analyze this voice data, it uses a speech recognition engine to perform recognition processing and convert it into text data. This process involves receiving the voice signal as digital data and converting it into text format via a speech recognition API.

[0795] Step 2:

[0796] The server analyzes the converted text data and retrieves the user's learning history and skill level from the database. The input is audio data in text format. Based on this data, the server dynamically selects appropriate learning questions using a generative AI model. It performs database queries and considers past performance and learning patterns to determine the most suitable quiz.

[0797] Step 3:

[0798] The server sends the selected problem as text data to the terminal. The input is the data of the training problems selected by the AI ​​model. The output is the data that the terminal uses speech synthesis technology to provide to the user with the selected problem in voice. This involves sending the data via the HTTP protocol and the synthesis process to convert it into speech.

[0799] Step 4:

[0800] The user responds to questions presented by the terminal using voice. The input is the user's own voice response. The terminal records this voice, converts it into a digital signal, and sends it back to the server. This is where the microphone-based voice capture and digital conversion functions come into play.

[0801] Step 5:

[0802] The server converts the voice response data received from the terminal back into text data using a speech recognition engine. The input is digitized voice data. The server analyzes this data and determines whether it is correct or incorrect. The generated text data is then passed to an evaluation algorithm, which includes an analysis process to determine whether the quiz answers are correct or incorrect.

[0803] Step 6:

[0804] The server generates a feedback message based on the judgment result and sends it to the terminal. The input is the correct / incorrect judgment result. The output is the feedback content to be provided to the user in voice. The server then converts this feedback into text and sends it to the terminal via communication.

[0805] Step 7:

[0806] The terminal plays feedback messages received from the server using an audio output device and conveys them to the user. The input is the text data of the feedback. The output is audio data that the user can hear. This includes the specific operation of converting text feedback into speech using speech synthesis technology and playing it back in real time.

[0807] (Application Example 1)

[0808] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0809] Modern industrial settings demand efficient worker training and safety assurance. However, traditional training methods struggle to provide individualized training, particularly in promoting understanding of machine operation and safety procedures among new employees. Maintaining learning motivation is also difficult. This can lead to decreased work efficiency and increased safety risks.

[0810] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0811] In this invention, the server includes means for the user to request access to educational content via voice input, means for analyzing the voice input and dynamically selecting appropriate questions based on the user's history and skill level, and means for presenting quizzes related to the operation and safety procedures of industrial machinery and educating workers using an audio assistant. This enables workers to receive personalized education and efficiently acquire machine operation and safety procedures.

[0812] "A means by which a user requests access to educational content via voice input" refers to a method for a user to request educational content using their voice.

[0813] "A means of analyzing voice input and dynamically selecting appropriate problems based on the user's history and skill level" refers to a method of analyzing voice data to select tasks in real time that correspond to the user's past learning record and skill level.

[0814] "Means of presenting selected questions to the user via an audio output device" refers to a method of presenting selected quizzes or tasks to the user via audio.

[0815] "A means of receiving a user's verbal response and converting it into corresponding text data using speech recognition technology" refers to a technology that receives a user's voice response and converts it into text.

[0816] "A means of evaluating text data, determining correctness, and generating immediate feedback" refers to a method of evaluating converted text, determining whether it is correct or incorrect, and providing rapid feedback.

[0817] "Means of providing generated feedback to the user in audio format" refers to methods of communicating evaluation results to the user in audio format.

[0818] "Means of recording user progress and analyzing data to improve future plans" refers to technologies that store learner performance and evaluate data to optimize future educational plans.

[0819] "A method of educating workers by presenting quizzes related to the operation and safety procedures of industrial machinery and using an audio assistant" refers to a method of training workers by presenting questions related to the handling and safety measures of industrial equipment and utilizing an audio assistant.

[0820] The system for realizing this application primarily consists of three components working together: a server, terminals (including robots), and the user. The server resides in the cloud and is equipped with a speech recognition engine and a generative AI model. The server receives voice input from the user, converts it to text, and then dynamically selects appropriate quiz questions, taking into account the user's learning history and skill level.

[0821] The terminal includes industrial robots and voice output devices, enabling voice-based interaction. When a user requests access to educational content by voice, the terminal presents the user with a quiz received from the server. The user's voice response is sent back to the server by the terminal, where it is judged as correct or incorrect, and feedback is provided to the user by voice.

[0822] Users can use this system to advance their workplace training. For example, when learning how to operate industrial machinery, the robot might ask a quiz question like, "What is the first step in the safety procedure for this machine?" If the user answers, "Turn off the power," the server will determine if the answer is correct, and the robot will immediately provide feedback such as, "That's correct. Next, let's check the emergency stop button."

[0823] This coordinated operation of the entire system enables efficient worker training. An example of a prompt message might be: "We are considering how to use this quiz assistant to educate workers on emergency stop procedures for safe machine operation. Please describe specific question examples and the feedback process."

[0824] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0825] Step 1:

[0826] The device receives voice input from the user. The device uses its built-in microphone to capture the user's voice request as a digital audio signal. This input is processed as a request to access educational content.

[0827] Step 2:

[0828] The device sends voice data to the server. The voice data is sent to the server as data packets over the internet. Based on this input, the server uses a speech recognition engine to convert the data into text data.

[0829] Step 3:

[0830] The server performs speech recognition. Using the Google Cloud Speech-to-Text API, it converts the speech data into text and generates user requests in text format. This output allows the server to understand what the user is requesting.

[0831] Step 4:

[0832] The server selects appropriate questions based on the user's history and skill level. It analyzes text data, queries the user's learning history database, and dynamically selects the most suitable quiz questions. This data processing ensures that appropriate educational content is prepared for the user.

[0833] Step 5:

[0834] The server sends a selected question to the terminal. The selected question is sent to the terminal as a data packet and presented to the user via an audio output device. This output allows the user to take on the quiz.

[0835] Step 6:

[0836] The user answers the quiz using their voice. The user answers the quiz using their voice, and this is captured again as an audio signal by the device.

[0837] Step 7:

[0838] The device sends the user's voice response to the server. The voice data is then sent back to the server in digital format and used as input for speech recognition technology to convert it into text data.

[0839] Step 8:

[0840] The server analyzes the user's response and determines whether it is correct or incorrect. The converted text is compared to the correct answer data for the question, and an AI model determines whether it is correct or not. Feedback is generated as the output of this data calculation.

[0841] Step 9:

[0842] The server generates feedback and sends it to the terminal. The generated feedback is sent to the terminal as voice instructions, providing the user with the results verbally. This output allows the user to know the accuracy of their answers.

[0843] Step 10:

[0844] Users receive feedback and prepare for the next challenge. Based on the feedback, they confirm their understanding and prepare to continue learning.

[0845] Through these steps, the system can provide users with efficient education and immediate feedback.

[0846] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0847] The present invention is a learning support system that incorporates an emotion engine to enable users to learn more effectively. This system monitors the user's learning progress in real time through interaction with the server, terminal, and user. In the process, it recognizes the user's emotional state and appropriately adjusts the learning content and feedback.

[0848] The server runs a program that integrates speech recognition technology and an emotion engine. First, the server processes the user's voice data received from the terminal and converts the content of the voice into text. During this process, it uses the emotion engine to infer the user's emotions from the tone and intonation of the voice. This emotion information is used to select learning content and adjust feedback. The server also uses the emotion data to generate content that reduces user stress and to create customized messages that increase motivation.

[0849] The device acts as an interface to provide learners with optimal quiz questions and feedback based on information obtained through emotion recognition. For example, if the device determines that the user is experiencing high stress levels, it can offer relaxing music or guidance. Furthermore, if the user shows positive emotions toward their learning goals, the device may present more challenging questions to encourage further effort.

[0850] Users participate in quizzes through their devices and answer by voice. During this process, the user's speech is sent to a server for speech recognition and sentiment analysis. For example, if a user appears bored while answering easy questions, the system will detect this through sentiment analysis and attempt to stimulate their motivation by providing new, more challenging questions. Furthermore, if a user is facing a difficult question, the device will play encouraging messages to support them.

[0851] This system allows users to receive appropriate learning tailored to their emotional state, enabling them to learn more effectively and efficiently. Furthermore, by utilizing an emotion engine, a more personalized learning experience is provided.

[0852] The following describes the processing flow.

[0853] Step 1:

[0854] The user gives a voice command to the device saying, "Start the quiz."

[0855] Step 2:

[0856] The terminal converts the user's voice commands into digital signals and sends an initiation request to the server.

[0857] Step 3:

[0858] The server receives requests from the terminal and checks the user's learning history and profile.

[0859] Step 4:

[0860] The server activates speech recognition and emotion engines, analyzing the user's speech data to infer their emotional state. This allows it to understand the user's current emotions.

[0861] Step 5:

[0862] The server selects the most suitable quiz questions from the database based on the user's learning history and emotional state.

[0863] Step 6:

[0864] The server sends the selected quiz questions to the terminal, allowing it to prepare them for presentation to the user.

[0865] Step 7:

[0866] The device converts the received quiz question into audio and asks the user, "Here's the next question: What is the capital of Japan?"

[0867] Step 8:

[0868] The user answers the quiz verbally, responding with "It's Tokyo."

[0869] Step 9:

[0870] The terminal receives the user's voice response and sends it to the server as a digital signal.

[0871] Step 10:

[0872] The server uses speech recognition technology to convert speech into text and evaluates the response by comparing it to correct answer data.

[0873] Step 11:

[0874] The server then performs sentiment analysis on the text data to analyze changes in emotions during the response.

[0875] Step 12:

[0876] The server generates customized feedback for the user based on the results of correct / incorrect judgment and sentiment analysis. For example, if the answer is correct, it might say, "That's correct, congratulations!" and if it's incorrect, it might say, "That's unfortunate, but let's try again next time."

[0877] Step 13:

[0878] The server provides the generated feedback to the user as audio output via the terminal.

[0879] Step 14:

[0880] The terminal communicates feedback from the server to the user via voice, helping to enhance the user's motivation to learn.

[0881] Through this series of processes, the system can grasp the user's emotional state in real time and highly personalize the learning experience.

[0882] (Example 2)

[0883] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0884] The problem that this invention aims to solve is to enable learners to receive an appropriate learning experience that is tailored to their emotional state. Conventional learning support systems have not taken into account the user's emotions when adjusting content, making it difficult to maximize learning efficiency. Furthermore, the learning content is uniform, and there is little feedback or content provision tailored to the individual interests and emotional state of the learners. As a result, it has been difficult to effectively maintain the learners' concentration and motivation.

[0885] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0886] This invention includes a server that analyzes voice input and dynamically selects appropriate problems based on the learner's past learning information and ability level; a server that uses emotion analysis technology to infer the user's emotions from the tone and intonation of the voice and adjusts the selection of learning materials and feedback; and a server that generates personalized messages for stress reduction and motivation enhancement based on the user's emotion data. This makes it possible to provide personalized learning content that reflects the user's emotions in real time.

[0887] "Voice input" refers to the act of a user communicating information or instructions to a system via voice.

[0888] A "learner" is someone who uses a system for the purpose of acquiring knowledge or skills.

[0889] "Past learning information" refers to records of learning content and results that learners have worked on in the past.

[0890] "Competency level" is an indicator that shows the degree of a learner's knowledge or skills.

[0891] "Dynamic selection" refers to making the optimal choice in real time based on the situation and conditions.

[0892] "Emotion analysis technology" is a technology that uses voice and other data to infer a user's emotional state.

[0893] "Voice tone and intonation" refers to the qualitative characteristics of a voice, such as pitch and intonation.

[0894] "Learning materials" refers to all learning content provided to learners.

[0895] "Feedback" refers to evaluations and guidance provided based on a learner's performance and progress.

[0896] "Stress reduction" refers to support aimed at alleviating the burden and tension of learners.

[0897] "Motivation enhancement" refers to methods and initiatives aimed at increasing learners' motivation.

[0898] A "personalized message" is a special message tailored to each user's situation and needs.

[0899] "Real-time" refers to a situation where processing and responses are performed instantly in response to the current circumstances.

[0900] "Personalized learning content" refers to learning materials optimized based on the individual characteristics and needs of each learner.

[0901] This invention is an advanced system for learning support that analyzes the user's emotional state in real time and provides personalized learning content accordingly.

[0902] The server first receives voice input from the user via the terminal. This is done using a microphone or voice input-enabled device. The received voice data is converted into text data using the Google Cloud Speech-to-Text API. Next, Affectiva's sentiment analysis technology is used to infer the user's emotions from the tone and intonation of the voice. These processes are carried out using computing resources within the data center.

[0903] The device's role is to provide users with optimal quiz questions and feedback based on emotional information and learning content provided by the server. For example, it can present application problems using mathematical formulas as quiz questions via audio. This allows users to progress through learning effectively.

[0904] Users respond to presented quizzes and feedback verbally. This audio is then sent back to the server for speech recognition and sentiment analysis.

[0905] As a concrete example, the following prompt can be entered into a generative AI model: "If the AI ​​detects heightened emotions from the voice of a user participating in the quiz, what kind of challenging question should be presented as the next learning content?"

[0906] This system allows users to find the learning method best suited to their emotional state, enabling them to learn more focused and efficiently than with traditional methods.

[0907] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0908] Step 1:

[0909] The server receives voice input from the user. It uses the Google Cloud Speech-to-Text API to convert the voice data, which is sent via the terminal, into text data. This conversion results in the voice data being output as text. Specifically, if the user says, "This is a simple quiz," the audio is reproduced exactly as it is in text.

[0910] Step 2:

[0911] The server uses Affectiva's sentiment analysis technology to infer the user's emotions based on the converted text data. It performs sentiment analysis based on the tone and intonation of the input voice and outputs the user's emotional state. Specifically, this includes a process where a bright tone of voice input is interpreted as "excited."

[0912] Step 3:

[0913] Based on the results of sentiment analysis, the server selects appropriate learning content, taking into account the user's past learning history and ability level. Using sentiment data and learning history data as input, it outputs optimal quiz questions and learning materials. For example, if the user is excited, it may select more difficult questions.

[0914] Step 4:

[0915] The terminal provides learning content presented by the server to the user. It prepares to present quiz questions and other content to the user via audio using its voice output function. The input is the learning content, and the audio output to the user is the output of this step. A concrete example would be a scenario where applied math problems are read aloud.

[0916] Step 5:

[0917] The user responds to the presented question using voice. They then input their voice again into the device, and this data is sent to the server. This serves as input for the next voice recognition step.

[0918] Step 6:

[0919] The server converts the user's voice response back into text data and performs a correctness check. It takes the response, converted into text by speech recognition, as input, determines whether it's correct, and generates feedback. This process outputs voice feedback such as "That's correct" for correct answers.

[0920] Step 7:

[0921] The server generates feedback and, taking into account the user's emotional state, devises additional support messages. The input is the previously obtained sentiment analysis data and correct / incorrect judgment results, while the output is a personalized message of encouragement or advice. For example, it might include a message like, "Let's try a more fun problem next time!"

[0922] Step 8:

[0923] The device delivers generated feedback and additional messages to the user via voice. The output of this step is delivered directly to the user's ears through voice output. Through these steps, the user receives a learning experience optimized for them.

[0924] (Application Example 2)

[0925] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0926] There is a need to address the issue of not being able to provide learners with appropriate feedback and learning experiences that take their emotional state into account when they use digital learning materials. This invention aims to improve learning efficiency and motivation by analyzing the learner's emotions through voice input and providing optimal learning content and feedback according to that state.

[0927] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0928] In this invention, the server includes means for the user to request access to digital learning materials through voice input, means for analyzing the voice input and dynamically selecting appropriate problems based on the learner's emotional state and knowledge level, and means for generating personalized messages to enhance learning motivation based on the user's emotional analysis. This enables a personalized learning experience that reflects the learner's emotional state.

[0929] An "educational plan" refers to a set of procedures and policies that take into account the goals and progress that learners should achieve, and that enable them to learn efficiently.

[0930] "Digital learning materials" refer to educational content provided using electronic media, and include learning materials in various formats such as text, video, and audio.

[0931] "Emotional state" refers to the learner's psychological and emotional condition and is a factor that influences the effectiveness of learning.

[0932] A "personalized learning experience" refers to a learning experience tailored to individual needs by providing learning content and feedback optimized according to each learner's characteristics and emotional state.

[0933] "Feedback" refers to information provided to learners that shows results and reactions, and is used to promote learning improvement and continuous growth.

[0934] The system implementing this invention consists of a server incorporating a digital learning support program using an emotion engine, and a terminal used by the user. First, the user makes a voice input through the terminal and requests access to digital learning materials. This request is sent to the server, which uses speech recognition technology to convert the voice into text information and performs emotion analysis.

[0935] The server infers the user's emotional state from the converted text information and dynamically generates questions and feedback tailored to the individual learner's characteristics based on the results. Using the emotion engine, if a learner is experiencing stress, it can present content to reduce stress or generate personalized messages to boost motivation.

[0936] The device presents feedback and learning content received from the server to the user via an audio output device. The user's verbal response is then sent back to the server, and feedback is provided immediately. This system allows the user to receive appropriate learning tailored to their emotional state.

[0937] For example, if a learner shows signs of anxiety while tackling a complex problem, the server will generate and output a message such as, "Don't rush, try again." An example of a prompt message for the generative AI model could be, "If the user is feeling anxious during learning, please provide an encouraging message."

[0938] The hardware includes a microphone and audio output device, while the software includes a speech recognition library and sentiment analysis algorithms. This combination makes it possible to provide learners with a personalized learning experience.

[0939] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0940] Step 1:

[0941] The user provides voice input to the device. The input voice data is captured via the device's microphone and converted into a digital signal. This data is then sent to the server.

[0942] Step 2:

[0943] The server uses speech recognition technology to convert the received audio data into text data. Using a speech recognition library, it analyzes the content of the audio and generates text as linguistic information. This text data then serves as input for sentiment analysis.

[0944] Step 3:

[0945] The server uses an emotion engine to analyze the user's emotional state from text data. This process employs an emotion analysis algorithm that outputs an emotion label (e.g., joy, sadness, boredom) based on the tone and wording of the text. This emotional information is then used in the next step.

[0946] Step 4:

[0947] The server selects the optimal learning content based on emotional information and the user's knowledge level. Here, a generative AI model is used to dynamically generate personalized questions and feedback according to the emotional labels. This generated content is output and sent to the terminal.

[0948] Step 5:

[0949] The terminal presents feedback and learning content received from the server to the user via an audio output device. The user verbally answers the presented questions, and this becomes the next input.

[0950] Step 6:

[0951] The user's verbal response is again captured as audio data on the device and sent to the server. The server uses speech recognition technology to convert the response into text and generates data to determine whether it is correct or incorrect.

[0952] Step 7:

[0953] The server evaluates the accuracy of the answer and immediately generates feedback. The generated feedback is then output back to the terminal and provided to the user. This allows the user to receive learning progress and recommendations in real time.

[0954] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0955] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0956] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0957] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0958] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0959] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0960] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0961] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0962] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0963] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0964] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0965] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0966] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0967] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0968] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0969] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0970] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0971] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0972] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0973] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0974] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0975] The following is further disclosed regarding the embodiments described above.

[0976] (Claim 1)

[0977] A means by which users request access to learning content through voice input,

[0978] A means for analyzing voice input and dynamically selecting appropriate quiz questions based on the learner's learning history and skill level,

[0979] A means of presenting selected quiz questions to the user via an audio output device,

[0980] A means for receiving a user's verbal response and converting it into corresponding text data using speech recognition technology,

[0981] A means of evaluating text data, determining its accuracy, and generating immediate feedback,

[0982] A means of providing the generated feedback to the user in voice,

[0983] A means of recording the user's learning progress and analyzing the data to improve future learning plans,

[0984] A system that includes this.

[0985] (Claim 2)

[0986] The system according to claim 1, wherein the speech recognition technology is optimized to correspond to the pronunciation and intonation of a specific language.

[0987] (Claim 3)

[0988] The system according to claim 1, which provides ranking and competition features based on quiz results in order to enhance user motivation for learning.

[0989] "Example 1"

[0990] (Claim 1)

[0991] A means by which the user requests access to learning resources through voice input,

[0992] A means for analyzing voice input and dynamically selecting appropriate questions based on the learner's history and skill level,

[0993] A means of presenting the selected problem to the user via an audio output device,

[0994] A means for receiving a user's verbal response and converting it into corresponding text data using recognition technology,

[0995] A means of evaluating text data, determining its accuracy, and instantly generating feedback,

[0996] A means of providing the generated feedback to the user in voice,

[0997] A means of recording the user's learning progress and analyzing the information to improve future plans,

[0998] A means of dynamically selecting problems and providing an individualized experience, taking into account learning history and patterns,

[0999] A system that includes this.

[1000] (Claim 2)

[1001] The system according to claim 1, wherein the recognition technology is optimized to correspond to the pronunciation and intonation of a specific language.

[1002] (Claim 3)

[1003] The system according to claim 1, which provides ranking and competition functions based on results in order to enhance motivation for learning.

[1004] "Application Example 1"

[1005] (Claim 1)

[1006] A means by which users request access to educational content through voice input,

[1007] A means for analyzing voice input and dynamically selecting appropriate problems based on the user's history and skill level,

[1008] A means of presenting the selected problem to the user via an audio output device,

[1009] A means for receiving a user's verbal response and converting it into corresponding text data using speech recognition technology,

[1010] A means of evaluating text data, determining its accuracy, and generating immediate feedback,

[1011] A means of providing the generated feedback to the user in voice,

[1012] A means of recording user progress and analyzing data to improve future plans,

[1013] A method of educating workers by presenting quizzes related to the operation and safety procedures of industrial machinery and using an audio assistant,

[1014] A system that includes this.

[1015] (Claim 2)

[1016] The system according to claim 1, wherein the speech recognition technology is optimized to support a specific language and accent.

[1017] (Claim 3)

[1018] The system according to claim 1, which provides ranking and competition functions based on quiz results to enhance user motivation and encourage competition among users.

[1019] "Example 2 of combining an emotion engine"

[1020] (Claim 1)

[1021] A means by which the user requests access to learning materials via voice input,

[1022] A means for analyzing voice input and dynamically selecting appropriate problems based on the learner's past learning information and ability level,

[1023] A means of presenting the selected problem to the user via an audio output device,

[1024] A means of receiving a user's verbal response and converting it into text information using speech recognition technology,

[1025] A means of evaluating text information, determining its accuracy, and generating immediate feedback,

[1026] A means of providing the generated feedback to the user in voice,

[1027] A means of recording the user's learning progress and analyzing the information to improve future learning plans,

[1028] A means of using emotion analysis technology to infer a user's emotions from their voice tone and intonation, and to adjust the selection of learning materials and feedback accordingly.

[1029] A means of generating personalized messages to reduce stress and improve motivation based on user emotional data,

[1030] A system that includes this.

[1031] (Claim 2)

[1032] The system according to claim 1, wherein the speech recognition technology is optimized to correspond to the pronunciation and intonation of a specific language, and the emotion analysis technology infers the user's emotional state based on the speech information.

[1033] (Claim 3)

[1034] The system according to claim 1, which provides ranking and competition features based on quiz results and sentiment information in order to enhance the user's motivation to learn, thereby increasing learning motivation.

[1035] "Application example 2 when combining with an emotional engine"

[1036] (Claim 1)

[1037] A means for users to request access to digital learning materials through voice input,

[1038] A means for analyzing voice input and dynamically selecting appropriate questions based on the learner's emotional state and knowledge level,

[1039] A means of presenting the selected problem to the user via an audio output device,

[1040] A means for receiving user speech and converting it into corresponding text information using speech recognition technology,

[1041] A means of evaluating textual information, determining its accuracy, and generating immediate feedback,

[1042] A means of providing the generated feedback to the user in voice,

[1043] A means for generating personalized messages to enhance learning motivation based on user sentiment analysis,

[1044] A means of recording users' learning progress and analyzing data to improve future educational plans,

[1045] A system that includes this.

[1046] (Claim 2)

[1047] The system according to claim 1, wherein the speech recognition technology corresponds to the pronunciation and intonation of a specific language, and an emotion estimation function is integrated.

[1048] (Claim 3)

[1049] The system according to claim 1, which enhances the user's motivation to learn and adjusts the difficulty level while taking into account their emotional state. [Explanation of Symbols]

[1050] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means by which users request access to learning content through voice input, A means for analyzing voice input and dynamically selecting appropriate quiz questions based on the learner's learning history and skill level, A means of presenting selected quiz questions to the user via an audio output device, A means for receiving a user's verbal response and converting it into corresponding text data using speech recognition technology, A means of evaluating text data, determining its accuracy, and generating immediate feedback, A means of providing the generated feedback to the user in voice, A means of recording the user's learning progress and analyzing the data to improve future learning plans, A system that includes this.

2. The system according to claim 1, wherein the speech recognition technology is optimized to correspond to the pronunciation and intonation of a specific language.

3. The system according to claim 1, which provides ranking and competition features based on quiz results in order to enhance the user's motivation to learn.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A