system

A system using facial expression and gesture analysis provides real-time feedback to instructors, addressing the challenge of delayed teaching method improvements by offering immediate and objective learner evaluation.

JP2026104367APending Publication Date: 2026-06-25SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-12-13
Publication Date
2026-06-25

AI Technical Summary

Technical Problem

Instructors face challenges in evaluating learners' understanding and concentration levels in real-time, leading to delayed and labor-intensive improvements in teaching methods, as traditional methods lack immediate and objective feedback.

Method used

A system utilizing information acquisition devices to collect learners' facial expressions and gestures, processed by an analysis device to evaluate comprehension and concentration levels, providing real-time notifications and generating specific suggestions for improving teaching methods.

Benefits of technology

Enables immediate adjustments to teaching methods based on real-time feedback, promoting instructor growth and improving educational quality by enhancing learner engagement and understanding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026104367000001_ABST
    Figure 2026104367000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] In an educational environment, means of acquiring information to analyze learners' level of understanding and concentration, An analytical means for evaluating the learner's emotional state and the effectiveness of the educator's explanation to the learner, using the information obtained by the aforementioned means, Based on the evaluation obtained by the aforementioned analysis means, a means for notifying educators in real time, A means for recording the aforementioned evaluation and notification content and generating suggestions for improving teaching methods, A means of analyzing the learner's level of understanding and concentration at home and notifying parents, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of this disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the modern educational environment, instructors are required to appropriately evaluate the understanding and concentration of learners under time, resource, and information constraints. In particular, it is difficult to grasp the reactions of learners in real time and take appropriate actions, and specific feedback for improving the quality of instruction is often not obtained. Also, with traditional methods, it is difficult to objectively and immediately evaluate the state of learners, and there is a problem that the time and labor required to improve the teaching method are large.

Means for Solving the Problems

[0005] This invention involves collecting information such as learners' facial expressions and gestures using an information acquisition device installed in the educational environment. The obtained information is processed by an analysis device to evaluate the learners' level of understanding and concentration. Based on this evaluation result, appropriate notifications are sent to instructors in real time, allowing instructors to immediately adjust their teaching methods. Furthermore, by providing a device that records the evaluation results and notification content and generates specific suggestions for improving teaching methods based on this information, the invention aims to promote instructor growth and improve the quality of education.

[0006] The term "educational environment" refers to the place and circumstances in which learning takes place, and includes all educational activities conducted within that environment.

[0007] A "learner" refers to a person who engages in activities to acquire knowledge and skills in an educational environment.

[0008] "Comprehension level" is an indicator that shows how well learners understand the content they have been taught.

[0009] "Concentration level" is an indicator that shows how much attention learners are paying to learning activities.

[0010] An "information acquisition device" refers to a device used to collect data on learners' facial expressions and actions within an educational environment.

[0011] An "analysis device" refers to a device that processes acquired data and is responsible for evaluating the learner's level of understanding and concentration.

[0012] The term "instructor" refers to a person who plays a role in guiding learners in knowledge and skills within an educational environment.

[0013] A "notification device" is a device that provides information to instructors in real time based on analysis results.

[0014] "Records" refer to information saved for use in later analysis and improvement, including evaluation results and notification content.

[0015] The "proposal generation device" is a device for specifically generating proposals regarding improvement of guidance methods based on the recorded information.

Brief Explanation of Drawings

[0016] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when the emotion engine is combined. [Figure 14]It is a sequence diagram showing the processing flow of a data processing system in Application Example 2 when a sentiment engine is combined.

Embodiment for Carrying out the Invention

[0017] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), etc.

[0020] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0021] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0022] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0024] [First Embodiment]

[0025] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0026] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0029] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0032] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0036] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0037] This invention relates to a system that analyzes learners' comprehension and concentration levels in an educational environment and provides real-time feedback to instructors. This system consists of an information acquisition device, an analysis device, a notification device, a recording device, and a suggestion generation device.

[0038] Information acquisition process

[0039] The server collects learners' facial expressions, posture, and voice data via cameras and microphones installed in the classroom. This provides foundational data to understand the learners' current emotional state. For example, the camera detects facial movements when a learner frowns and records this as a change in facial expression.

[0040] Data analysis process

[0041] The server processes the collected data using advanced image and audio analysis algorithms. Image data is analyzed using facial recognition technology to determine whether the learner understands or is confused. Audio data is analyzed using audio analysis technology to extract characteristics of the instructor's speaking style (e.g., intonation and tempo) and evaluate the effectiveness of the instruction.

[0042] Real-time notification process

[0043] The device notifies the instructor in real time if it detects a decrease in the learner's concentration or that they are having difficulty understanding the material. For example, a message such as "Student D may be having difficulty understanding" might appear on the instructor's device. This information can then be used by the instructor to adjust their teaching methods on the spot.

[0044] Report generation process

[0045] After the lesson ends, the server automatically generates a comprehensive evaluation report based on the data collected. The report specifically shows which learners made progress in understanding and at what point their concentration wavered. This allows instructors to review the entire lesson and use the results to improve their teaching later.

[0046] Proposal generation process

[0047] The server is a device that uses recorded data and analysis results to generate improvement suggestions for instructors. This allows instructors to obtain specific, actionable directions for further deepening their teaching methods. For example, it might provide suggestions such as, "Incorporating visual aids into lessons may improve comprehension."

[0048] This system enables the integrated and dynamic assessment and response to learners' understanding and concentration within the educational environment, thereby promoting improvements in the quality of education.

[0049] The following describes the processing flow.

[0050] Step 1:

[0051] The server collects video and audio data in real time through cameras and microphones installed in the classroom. The cameras capture each student's face and expressions, and the microphones record the instructor's voice. During this process, the image data is processed frame by frame, and the audio data is converted to the appropriate format.

[0052] Step 2:

[0053] The server preprocesses the collected image and audio data. This preprocessing involves performing face recognition and expression extraction using the image data, and removing noise from the audio data to generate clear audio clips. This process prepares the data for analysis.

[0054] Step 3:

[0055] The server inputs preprocessed data into an AI model and performs analysis using a deep learning algorithm. The analysis classifies the learner's emotional state from their facial expressions and evaluates their concentration and comprehension levels. It also analyzes the instructor's speaking speed and intonation from audio data to determine if the instruction is effective.

[0056] Step 4:

[0057] The device receives analysis results from the server and sends real-time notifications to the instructor. These notifications include alerts if the learner's comprehension is declining or if their speaking style needs improvement. These notifications appear as pop-ups on the instructor's device.

[0058] Step 5:

[0059] At the end of each lesson, the server generates a report based on an overall assessment of comprehension and concentration levels derived from analysis. This report includes each learner's responses and changes in concentration at specific points during the lesson. Instructors can use this information to improve future lessons.

[0060] Step 6:

[0061] The server generates specific suggestions for improving teaching methods based on accumulated data. These suggestions include concrete actions to improve learner responses based on past data, supporting the growth of instructors.

[0062] (Example 1)

[0063] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0064] In educational settings, there is a lack of means to accurately grasp learners' levels of understanding and concentration in real time and to quickly convey that information to instructors. Furthermore, there is a need for effective methods to appropriately analyze changes in learners' emotional states and learning progress and use that information to improve instruction. This presents a challenge in that it is difficult for instructors to adjust their teaching methods to be optimal for each learner in a timely manner.

[0065] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0066] In this invention, the server includes means for acquiring the learner's emotional state, posture, and voice information in the learning environment; means for analyzing the learner's emotional state and the characteristics of the instructor's explanation using the image and voice information acquired by the means; and means for providing notifications that enable real-time adjustment of the teaching method based on the analysis results. This makes it possible to grasp the learner's level of understanding and concentration in real time and dynamically adjust the teaching method based on that.

[0067] The term "learning environment" refers to the place or space where education takes place, as well as the equipment and technology used within it.

[0068] A "learner" refers to an individual who participates in educational activities with the aim of acquiring knowledge and skills.

[0069] "Emotional state" refers to the psychological expressions and emotional states exhibited by learners, which can be observed through changes in facial expressions and voice.

[0070] "Postural information" refers to data related to the learner's body position and movement, and is used to determine their level of concentration and interest.

[0071] "Auditory information" refers to sound data collected within the learning environment, encompassing the speech of learners and instructors, as well as ambient sounds.

[0072] "Analysis results" refer to the conclusions and insights derived from processing the collected data.

[0073] "Notifications" refer to information and messages generated based on analysis and provided to instructors.

[0074] "Teaching methods" refer to the techniques and processes by which an instructor provides education to learners.

[0075] "Recording" refers to the act of saving data and information so that it can be referenced or analyzed later, or the result of doing so.

[0076] "Improvement suggestions" refer to specific proposals provided to improve the quality of education, based on recorded data and analysis.

[0077] This invention aims to implement a system that analyzes learners' comprehension and concentration levels in real time within educational settings and provides feedback to instructors. The details are described below.

[0078] The server utilizes cameras and microphones installed within the learning environment to collect learners' facial expressions, posture, and voice information. This data forms the basis for understanding the learners' emotional state at the time of acquisition.

[0079] Next, the server processes the collected data using advanced image and audio analysis software. Image data is analyzed based on facial expression recognition technology to identify how well the learner understands or is confused. Meanwhile, audio data is analyzed for changes in tone and tempo to evaluate the characteristics of the instructor's explanation.

[0080] The terminal notifies the instructor in real time based on the judgment results obtained from the server. This notification is displayed to the instructor in the form of, for example, "Student D may be having difficulty understanding." The instructor uses this information to adjust the lesson plan and approach on the spot.

[0081] Furthermore, after the lesson ends, the server automatically generates an evaluation report based on the analysis results. This report details which learners showed difficulties with understanding or concentration and at what point, and is used as feedback for instructors.

[0082] Furthermore, based on the recorded data and evaluation results, suggestions for improving teaching methods are generated. For example, a suggestion might be made such as, "Incorporating visual materials into lessons may improve learners' comprehension."

[0083] This system will promote increased efficiency and improved quality of instruction in educational settings.

[0084] Examples of prompt statements include the following:

[0085] "Could you provide an example of a notification message sent to an instructor if they determine that a student is losing focus during class?"

[0086] "Please provide three suggestions for improving learners' comprehension."

[0087] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0088] Step 1:

[0089] The server acquires the learner's facial expressions, posture, and voice information via the camera and microphone within the learning environment. It receives camera video and audio data as input and uses this to aggregate basic data about the learner's state in real time. Specifically, the image sensor captures facial expression data, and the voice sensor records the learner's speech and ambient sounds. The collected raw data is sent to the analysis module as output.

[0090] Step 2:

[0091] The server performs facial recognition using collected image data. It receives raw data as input and applies image processing algorithms to classify the learner's emotional state. For example, it determines states such as "understanding," "confused," and "indifferent" from eyebrow movements and changes in the mouth. As output, these analysis results are quantified as levels of concentration and comprehension and stored in a database.

[0092] Step 3:

[0093] The server performs speech analysis using audio data. It receives recorded audio as input and extracts speech features by applying a speech recognition algorithm. Specifically, it analyzes changes in the instructor's speaking tempo and volume and compares them with the learner's reactions. As output, an evaluation of the effectiveness of the instruction content is quantified and integrated with the results of facial expression analysis.

[0094] Step 4:

[0095] The device generates real-time notifications based on the analysis results. It receives analysis data as input and analyzes and makes decisions based on set thresholds. For example, if the "concentration level" falls below a predetermined value, it notifies the instructor that "Person A's concentration is declining." Actionable feedback is displayed on the device screen as output.

[0096] Step 5:

[0097] The server automatically generates evaluation reports based on the integrated data after each lesson. It receives the aggregated analytical data as input and processes it to visualize the progress of each learner's understanding and concentration levels. Specifically, the report outputs areas where understanding has improved and areas requiring improvement. The output is provided in a format viewable by instructors, which can be used to improve future lessons.

[0098] Step 6:

[0099] The server generates specific suggestions for improving teaching methods based on recorded data. Using analysis results stored in the database as input, it employs an AI model to derive optimal improvement suggestions. The output provides instructors with suggestions such as "increase the use of visual aids," which can be used to plan future lessons.

[0100] (Application Example 1)

[0101] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0102] In educational settings and at home, there is a growing need to dynamically grasp learners' comprehension and concentration levels and provide real-time feedback. However, traditional methods have faced challenges such as delayed feedback in educational settings and a lack of data necessary for providing appropriate guidance to individual learners. Furthermore, at home, it is currently difficult for parents to accurately grasp their children's learning progress and take appropriate action.

[0103] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0104] In this invention, the server includes means for acquiring information for analyzing the learner's level of understanding and concentration in an educational environment; analysis means for evaluating the learner's emotional state and the effectiveness of the educator's explanations to the learner using the information acquired by the means; and means for notifying the educator in real time based on the evaluation obtained by the analysis means. This enables the learner's condition to be grasped quickly and accurately, allowing educators and parents to take immediate action.

[0105] "Educational environment" is a general term for the places and circumstances in which educational activities are conducted for learners.

[0106] A "learner" refers to someone who receives education in order to acquire knowledge and skills.

[0107] "Comprehension level" is a measure that indicates the extent to which a learner understands the learning material.

[0108] "Concentration level" is a measure that indicates how much attention a learner is paying to a learning activity.

[0109] "Means of acquiring information" refers to devices and methods for collecting data about learners.

[0110] "Analysis means" refers to the processes and techniques used to analyze acquired information and evaluate the learner's state.

[0111] "Means of notification" refers to methods or devices for conveying specific information to educators or parents based on analysis results.

[0112] "Guardian" refers to an adult who is responsible for the upbringing of a learner.

[0113] "Means of generating suggestions" refers to methods and processes for creating suggestions that help improve teaching methods based on acquired and analyzed data.

[0114] This system aims to analyze learners' comprehension and concentration levels in educational and home environments and provide real-time feedback to educators and parents. The following describes its configuration and operating procedures in detail.

[0115] The server uses cameras and microphones installed in educational and home environments to acquire learners' facial expressions, posture, and voice data. This allows for the collection of data tailored to the situation in real time. The hardware uses a standard webcam as the camera and a directional microphone as the microphone.

[0116] The collected data is processed on a server. For image data, facial recognition is performed using OpenCV and TENSORFLOW®. Audio data is analyzed using PyDub to evaluate intonation and tempo. This analysis allows us to determine whether the learner understands or is confused.

[0117] Based on the analysis results, the device sends notifications to educators and parents. Using the Pushbullet API, messages are sent in real time to smartphones and tablets. For example, a notification might appear stating, "Learner A may be having difficulty understanding the material." This information can be used to immediately adjust teaching methods.

[0118] Furthermore, after a learning session ends, the server automatically generates a detailed report based on the collected and analyzed data. This report is provided to educators as a reference for improving their teaching methods and can also be used to support learning at home.

[0119] As a concrete example, when a student is working on a math assignment at home, facial recognition can detect signs of confusion, such as wrinkles between the eyebrows or changes in eye movement. The parent is then immediately notified that "Student A appears to be having difficulty understanding the material."

[0120] An example of a prompt to input into the generating AI model is: "Analyze my child's facial expressions while they are studying math, determine if they are having difficulty understanding, and send a notification."

[0121] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0122] Step 1:

[0123] The server collects learners' facial expressions, posture, and voice data in real time through installed cameras and microphones. This input data includes changes in facial expressions and vocal intonation, and is used as foundational data to understand the learners' emotional state. Specifically, the camera captures facial feature points, and the microphone records the tone and volume of the voice.

[0124] Step 2:

[0125] The server performs facial recognition processing on the acquired image data using OpenCV and TensorFlow. This process extracts feature points from the learner's face from the input data and analyzes what emotions the input facial image indicates. For example, if there are wrinkles between the eyebrows, it will be determined that the person is confused. The output will be the analysis result of the learner's emotional state.

[0126] Step 3:

[0127] The server uses PyDub to perform speech analysis on the audio data. This analysis examines the intonation and tempo of the input audio data to determine the learner's level of concentration and comprehension. Specifically, if the tone of voice suddenly becomes flat, it is determined that the learner's concentration may have been interrupted. The output of this analysis represents the results of the comprehension and concentration assessment.

[0128] Step 4:

[0129] Based on the analysis results, the device uses the Pushbullet API to send notifications to educators and parents. In this step, the device displays the analyzed results of the learner's comprehension and concentration levels as a message. For example, a notification saying "Learner A may be having difficulty understanding" is sent to the parent's smartphone. The input is the analysis results, and the output is the notification message.

[0130] Step 5:

[0131] The server automatically generates a comprehensive evaluation report based on the data collected after each learning session. This report includes information on when the learner understood the material or when their concentration wavered, providing valuable insights for improving future teaching methods. Specifically, the server organizes the collected data chronologically and summarizes the evaluation results. The input is the previously analyzed data, and the output is the evaluation report.

[0132] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0133] This invention combines an emotion engine with a system for analyzing learners' comprehension and concentration levels in an educational environment. This system consists of an information acquisition device, an analysis device, a notification device, a recording device, a suggestion generation device, and an emotion engine.

[0134] The process of information acquisition and emotion recognition

[0135] The server collects learners' facial expressions and voice data in real time through cameras and microphones installed in the classroom. The emotion engine uses this data to recognize the learners' emotional state and identify emotions such as "interesting" or "confusing."

[0136] Data analysis and coaching feedback process

[0137] The server inputs data, including emotional states recognized by the emotion engine, into the analysis device. The analysis device evaluates the relationship between the learner's emotional state, comprehension level, and concentration level, and makes a comprehensive judgment on the effectiveness of the instruction. As a result, it is possible to understand how the learner is receiving the lesson content. For example, if a learner is "confused," it is determined that the teaching method at that point should be reconsidered.

[0138] Real-time notifications and the process of improving instruction

[0139] The device provides real-time feedback to instructors based on the analysis results. For example, a message such as "Ms. E seems a little confused" might appear on the instructor's device, allowing the instructor to adjust their teaching method immediately. Feedback is also provided if certain emotions are repeatedly observed.

[0140] Recording and proposal generation process

[0141] The server accumulates past emotional data and analysis results to evaluate changes in learners' states over the long term. Based on this, it generates suggestions for optimizing teaching methods. For example, if it is confirmed that some learners' comprehension improves with audiovisual stimuli, the use of visual materials will be suggested.

[0142] This system enables the improvement of educational quality by conducting multifaceted analysis, including learners' emotions. Instructors can grasp learners' emotions in real time and gain concrete means to implement appropriate instruction.

[0143] The following describes the processing flow.

[0144] Step 1:

[0145] The server uses cameras and microphones in the classroom to collect learners' facial expressions and audio data in real time. The cameras capture each learner's facial expressions, and the microphones record the tone and intonation of their individual voices. This data is temporarily stored in a database in its raw state.

[0146] Step 2:

[0147] The server preprocesses the collected data. Image data undergoes image filtering to identify facial features and extract expressions. Audio data is processed through noise filtering and audio frame extraction, and then formatted to be suitable for audio analysis methods.

[0148] Step 3:

[0149] The server supplies pre-processed data to the emotion engine. The emotion engine uses machine learning algorithms to analyze emotions from facial expressions and voice, determining emotional states such as "joy," "confusion," and "distraction." This enables instantaneous diagnosis of the learner's emotions.

[0150] Step 4:

[0151] The device receives sentiment analysis results sent from the server and notifies the instructor. Specifically, it displays a message such as, "Student F is confused by the lesson content," prompting the instructor to adjust the lesson's pace and explanation methods. This notification enables flexible responses during the lesson.

[0152] Step 5:

[0153] The server tracks and stores emotional data and associated analytical information over long periods. The stored data is organized in a format that allows for the evaluation of emotional changes over time. As a result, instructors can understand the temporal evolution of learners' emotions and use this information to improve teaching methods.

[0154] Step 6:

[0155] The server generates specific suggestions for optimizing teaching methods based on the combined data. These suggestions include actionable improvements, such as "Increasing visual aids can improve comprehension," providing practical guidelines for instructors.

[0156] (Example 2)

[0157] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0158] In educational settings, it is difficult to instantly grasp learners' levels of understanding and concentration and provide appropriate instruction; therefore, efficient methods are needed to maximize learning effectiveness. In particular, there is a need to develop a system that allows instructors to respond quickly by providing real-time feedback on learners' emotional states.

[0159] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0160] In this invention, the server includes means for collecting video and audio from the environment to identify the learner's emotions; means for analyzing the collected video and audio and evaluating the learner's emotions and learning effectiveness using a specified algorithm; and means for providing real-time feedback to the instructor according to the evaluated emotional state. This makes it possible to instantly grasp the learner's level of understanding and emotional state and improve the quality of instruction.

[0161] A "learner" is someone who seeks to acquire knowledge and skills in an educational environment.

[0162] "Emotions" refer to the mental state of a learner and include states such as interest, confusion, happiness, and anxiety.

[0163] "Video" refers to dynamic or static image data that visually represents the learner's facial expressions and actions.

[0164] "Speech" refers to acoustic signals, including the learner's utterances and tone of voice.

[0165] An "algorithm" is a set of procedures or computational steps for solving a specific problem, and it is used as a means of analyzing learners' emotions.

[0166] "Feedback" is information provided to instructors regarding the learner's current state, which allows them to adjust their teaching methods accordingly.

[0167] A "threshold" is a specific numerical value or level set as a standard for evaluating a learner's level of understanding and concentration.

[0168] A "suggestion" is information that provides instructors with specific guidance on improving and optimizing teaching methods, based on recorded data and analysis results.

[0169] This invention is a system for analyzing learners' comprehension and concentration levels in an educational environment and optimizing teaching methods. The system consists of three main components: a server, terminals, and users.

[0170] First, the server uses cameras and microphones placed in the classroom to collect learners' facial expressions and voice data in real time. This hardware operates continuously to facilitate data acquisition. The collected data is processed by an emotion engine to analyze the learners' emotional state. This emotion engine is equipped with emotion recognition algorithms and has advanced capabilities to analyze subtle changes in facial expressions and tone of voice.

[0171] Next, based on this analyzed data, the device provides real-time feedback to the instructor. The device immediately notifies the instructor according to the analysis results, displaying specific messages such as, "Student E is confused." This allows the instructor to immediately adopt teaching methods that respond to the learner's situation.

[0172] Furthermore, the server accumulates this data and analysis results over the long term, continuously evaluating changes in learners' situations. Based on this long-term data, the server can generate specific suggestions for improving teaching methods. This allows instructors to obtain valuable information for continuously optimizing their teaching methods.

[0173] For example, if data reveals that visual learning materials improve the comprehension of certain learners, then teaching methods can be proposed. Instructors can then use these suggestions to further improve the quality of their instruction.

[0174] An example of a prompt for a generative AI model is: "Please explain how to analyze learners' emotions in a classroom and evaluate their comprehension and concentration levels. Also, show how this data can be used to optimize instruction."

[0175] This system enables data-driven instruction to improve the quality of education.

[0176] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0177] Step 1:

[0178] The server uses cameras and microphones installed in the classroom to collect learners' facial expressions and voice data as input. Specifically, the cameras capture video at several frames per second and send it to the server as digital data. The microphones record learners' speech and ambient sounds in real time and send them to the server in digital format. This generates a petabyte-scale dataset that reflects the learners' state.

[0179] Step 2:

[0180] The server takes the collected facial expression and audio data as input and entrusts the processing to the emotion engine. Specifically, the emotion engine uses a facial recognition algorithm to analyze the facial features of each frame and uses speech recognition technology to evaluate the tone of voice and emotion. As output, the server generates emotion states labeled as "interesting" or "confused." These labels indicate the learner's state in real time.

[0181] Step 3:

[0182] The server sends emotional state data obtained from the emotion engine as input to the analysis device. The analysis device uses data mining techniques to compare this data with past data and evaluate the learner's level of understanding and concentration. As output, evaluation results regarding the effectiveness of specific teaching methods are generated. This prepares the server to receive feedback on the effectiveness of the instruction as numerical values ​​and evaluation comments.

[0183] Step 4:

[0184] The terminal receives evaluation results from the server as input and provides real-time feedback to the instructor. Specifically, the terminal displays a warning message such as "Person E is confused," allowing the instructor to immediately adjust their teaching methods based on the feedback. The output is information provided to the instructor and improvements to the teaching based on that information. This rapid feedback enables instructors to respond immediately to provide appropriate instruction.

[0185] Step 5:

[0186] The server records all data and analysis results over the long term as input. Specifically, it stores data in a database and organizes it to be useful for subsequent analysis. Based on this data, a suggestion generator generates suggestions for optimizing teaching methods. As output, instructors are provided with concrete suggestions for long-term educational improvement. This establishes a sustainable improvement cycle for improving the quality of education.

[0187] (Application Example 2)

[0188] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0189] In educational settings, it is difficult to appropriately assess learners' comprehension and concentration levels and provide real-time feedback. Traditional systems fail to adequately grasp learners' emotional states, making it difficult to improve teaching methods to suit individual learners. Furthermore, while there is a need to respond quickly to situations where learners are confused and provide appropriate educational materials, there is a lack of effective means to achieve this.

[0190] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0191] In this invention, the server includes means for acquiring information to analyze learners' awareness and concentration levels in an educational environment, means for evaluating learners' emotional states and the effectiveness of instructors' explanations using the acquired information, and means for notifying instructors in real time based on the evaluation and providing interactive educational materials to promote understanding among learners. This makes it possible to instantly grasp learners' emotional states and provide optimal teaching methods and materials for each individual learner.

[0192] The term "educational environment" refers to the place or situation in which learners acquire knowledge and skills, and includes physical classrooms and digital platforms.

[0193] The term "learner" refers to individuals who seek to acquire knowledge and skills within an educational environment.

[0194] "Recognition level" is an indicator that shows how well learners understand the educational content.

[0195] "Concentration" refers to the state in which learners are paying attention to educational activities.

[0196] "Emotional state" refers to the psychological state or reaction that learners exhibit during learning, and includes specific emotions such as "interesting" or "confusing."

[0197] "Interactive educational materials" refer to learning content and materials that allow learners to actively participate and interact with each other.

[0198] "Evaluation" refers to the process of analyzing learners' comprehension, concentration levels, and emotional states based on acquired data, and measuring the effectiveness of teaching methods.

[0199] "Notification" refers to the act of communicating information to instructors based on analysis results.

[0200] "Feedback" refers to information provided to improve teaching methods and enhance educational effectiveness, based on evaluation results, learner responses, and other factors.

[0201] The system implementing this invention analyzes learners' emotional states and comprehension levels in real time within an educational environment and provides appropriate feedback. The server uses cameras and microphones installed in the educational space to collect learners' facial expressions and voices in real time. This collected data is analyzed through emotion recognition software, such as Microsoft® Azure® Emotion API, to detect the learners' emotional states.

[0202] The device receives analysis results from the server and notifies the instructor in real time with feedback tailored to the learner's emotional state. This notification includes specific advice for the instructor to adjust their teaching methods according to the learner's condition.

[0203] Furthermore, the server records changes in learners' comprehension and emotional states, and generates suggestions for evaluating long-term learning effectiveness. These suggestions include guidelines on how interactive educational materials should be used.

[0204] For example, in a physics lesson in a classroom, the server might detect that a student is experiencing "confusion" when faced with a difficult problem. The terminal would immediately notify the instructor with a message such as, "Some students are confused. Please try explaining using concrete examples." This notification allows the instructor to flexibly adjust their teaching methods.

[0205] Possible prompt statements for input to a generative AI model include the following:

[0206] "How can you engage students who aren't interested in historical discussions?"

[0207] "When we detect that students are confused while solving math problems, what kind of interactive learning materials can we provide?"

[0208] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0209] Step 1:

[0210] The server collects learners' facial expressions and audio data in real time through cameras and microphones installed in the educational space. Inputs include video streams and audio streams, which are stored in a database. The operations performed at this stage involve video capture by the cameras and audio recording by the microphones.

[0211] Step 2:

[0212] The server sends the collected facial expression and voice data to emotion recognition software (e.g., Microsoft Azure Emotion API) to analyze the learner's emotional state. The input is in the form of raw data, which the emotion recognition software processes to generate outputs representing the learner's intuitive psychological state, such as labels like "interesting" or "confused." Data preprocessing and emotion analysis are the main operations in this step.

[0213] Step 3:

[0214] The server inputs emotional state data obtained from emotion recognition software into an analysis device to evaluate learners' comprehension and concentration levels. At this stage, data calculations are performed to correlate emotional states with the progress of the lesson and generate evaluation results. The output provides individual learners' evaluations of comprehension and concentration levels.

[0215] Step 4:

[0216] The terminal notifies instructors of evaluation results in real time, helping them select appropriate teaching methods. The input is the evaluation results sent from the server, and the output is a specific feedback message presented to the instructor. The key feature here is that the notification system operates in real time.

[0217] Step 5:

[0218] The server stores evaluation results and emotional state records in a database and generates suggestions for improving teaching methods based on that data. Long-term learner data history is used as input, and actionable suggestions are generated as output. The main operations in this step are data analysis and suggestion generation.

[0219] Step 6:

[0220] The terminal provides the instructor with generated suggestions and directly offers learners interactive educational materials tailored to specific situations as needed. Input is data from the suggestion generator, and output is specific educational material. The operation here is an appropriate response based on the learner's situation.

[0221] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0222] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0223] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0224] [Second Embodiment]

[0225] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0226] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0227] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0228] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0229] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0230] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0231] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0232] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0233] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0234] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0235] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0236] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0237] This invention relates to a system that analyzes learners' comprehension and concentration levels in an educational environment and provides real-time feedback to instructors. This system consists of an information acquisition device, an analysis device, a notification device, a recording device, and a suggestion generation device.

[0238] Information acquisition process

[0239] The server collects learners' facial expressions, posture, and voice data via cameras and microphones installed in the classroom. This provides foundational data to understand the learners' current emotional state. For example, the camera detects facial movements when a learner frowns and records this as a change in facial expression.

[0240] Data analysis process

[0241] The server processes the collected data using advanced image and audio analysis algorithms. Image data is analyzed using facial recognition technology to determine whether the learner understands or is confused. Audio data is analyzed using audio analysis technology to extract characteristics of the instructor's speaking style (e.g., intonation and tempo) and evaluate the effectiveness of the instruction.

[0242] Real-time notification process

[0243] The device notifies the instructor in real time if it detects a decrease in the learner's concentration or that they are having difficulty understanding the material. For example, a message such as "Student D may be having difficulty understanding" might appear on the instructor's device. This information can then be used by the instructor to adjust their teaching methods on the spot.

[0244] Report generation process

[0245] After the lesson ends, the server automatically generates a comprehensive evaluation report based on the data collected. The report specifically shows which learners made progress in understanding and at what point their concentration wavered. This allows instructors to review the entire lesson and use the results to improve their teaching later.

[0246] Proposal generation process

[0247] The server is a device that uses recorded data and analysis results to generate improvement suggestions for instructors. This allows instructors to obtain specific, actionable directions for further deepening their teaching methods. For example, it might provide suggestions such as, "Incorporating visual aids into lessons may improve comprehension."

[0248] This system enables the integrated and dynamic assessment and response to learners' understanding and concentration within the educational environment, thereby promoting improvements in the quality of education.

[0249] The following describes the processing flow.

[0250] Step 1:

[0251] The server collects video and audio data in real time through cameras and microphones installed in the classroom. The cameras capture each student's face and expressions, and the microphones record the instructor's voice. During this process, the image data is processed frame by frame, and the audio data is converted to the appropriate format.

[0252] Step 2:

[0253] The server preprocesses the collected image and audio data. This preprocessing involves performing face recognition and expression extraction using the image data, and removing noise from the audio data to generate clear audio clips. This process prepares the data for analysis.

[0254] Step 3:

[0255] The server inputs preprocessed data into an AI model and performs analysis using a deep learning algorithm. The analysis classifies the learner's emotional state from their facial expressions and evaluates their concentration and comprehension levels. It also analyzes the instructor's speaking speed and intonation from audio data to determine if the instruction is effective.

[0256] Step 4:

[0257] The device receives analysis results from the server and sends real-time notifications to the instructor. These notifications include alerts if the learner's comprehension is declining or if their speaking style needs improvement. These notifications appear as pop-ups on the instructor's device.

[0258] Step 5:

[0259] At the end of each lesson, the server generates a report based on an overall assessment of comprehension and concentration levels derived from analysis. This report includes each learner's responses and changes in concentration at specific points during the lesson. Instructors can use this information to improve future lessons.

[0260] Step 6:

[0261] The server generates specific suggestions for improving teaching methods based on accumulated data. These suggestions include concrete actions to improve learner responses based on past data, supporting the growth of instructors.

[0262] (Example 1)

[0263] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0264] In educational settings, there is a lack of means to accurately grasp learners' levels of understanding and concentration in real time and to quickly convey that information to instructors. Furthermore, there is a need for effective methods to appropriately analyze changes in learners' emotional states and learning progress and use that information to improve instruction. This presents a challenge in that it is difficult for instructors to adjust their teaching methods to be optimal for each learner in a timely manner.

[0265] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0266] In this invention, the server includes means for acquiring the learner's emotional state, posture, and voice information in the learning environment; means for analyzing the learner's emotional state and the characteristics of the instructor's explanation using the image and voice information acquired by the means; and means for providing notifications that enable real-time adjustment of the teaching method based on the analysis results. This makes it possible to grasp the learner's level of understanding and concentration in real time and dynamically adjust the teaching method based on that.

[0267] The term "learning environment" refers to the place or space where education takes place, as well as the equipment and technology used within it.

[0268] A "learner" refers to an individual who participates in educational activities with the aim of acquiring knowledge and skills.

[0269] "Emotional state" refers to the psychological expressions and emotional states exhibited by learners, which can be observed through changes in facial expressions and voice.

[0270] "Postural information" refers to data related to the learner's body position and movement, and is used to determine their level of concentration and interest.

[0271] "Auditory information" refers to sound data collected within the learning environment, encompassing the speech of learners and instructors, as well as ambient sounds.

[0272] "Analysis results" refer to the conclusions and insights derived from processing the collected data.

[0273] "Notifications" refer to information and messages generated based on analysis and provided to instructors.

[0274] "Teaching methods" refer to the techniques and processes by which an instructor provides education to learners.

[0275] "Recording" refers to the act of saving data and information so that it can be referenced or analyzed later, or the result of doing so.

[0276] "Improvement suggestions" refer to specific proposals provided to improve the quality of education, based on recorded data and analysis.

[0277] This invention aims to implement a system that analyzes learners' comprehension and concentration levels in real time within educational settings and provides feedback to instructors. The details are described below.

[0278] The server utilizes cameras and microphones installed within the learning environment to collect learners' facial expressions, posture, and voice information. This data forms the basis for understanding the learners' emotional state at the time of acquisition.

[0279] Next, the server processes the collected data using advanced image and audio analysis software. Image data is analyzed based on facial expression recognition technology to identify how well the learner understands or is confused. Meanwhile, audio data is analyzed for changes in tone and tempo to evaluate the characteristics of the instructor's explanation.

[0280] The terminal notifies the instructor in real time based on the judgment results obtained from the server. This notification is displayed to the instructor in the form of, for example, "Student D may be having difficulty understanding." The instructor uses this information to adjust the lesson plan and approach on the spot.

[0281] Furthermore, after the lesson ends, the server automatically generates an evaluation report based on the analysis results. This report details which learners showed difficulties with understanding or concentration and at what point, and is used as feedback for instructors.

[0282] Furthermore, based on the recorded data and evaluation results, suggestions for improving teaching methods are generated. For example, a suggestion might be made such as, "Incorporating visual materials into lessons may improve learners' comprehension."

[0283] As a result, this system promotes the efficiency and quality improvement of teaching in educational settings.

[0284] Specific examples of prompt sentences include the following:

[0285] "When it is determined that a student has lost concentration during class, please provide an example of a notification message sent to the instructor."

[0286] "Please provide three suggestions for improving the learner's understanding."

[0287] The flow of the specific process in Example 1 will be described using FIG. 11.

[0288] Step 1:

[0289] The server acquires the learner's facial expressions, postures, and voice information via cameras and microphones within the learning environment. It receives camera images and voice data as input and aggregates basic data regarding the learner's state in real time based on this. As a specific operation, the image sensor captures facial expression data, and the voice sensor records the learner's speech and ambient sounds. As output, the collected raw data is sent to the analysis module.

[0290] Step 2:

[0291] The server performs facial expression recognition using the collected image data. It receives raw data as input and applies image processing algorithms to classify the learner's emotional state. For example, it determines states such as "understanding," "confusion," "indifference," etc. from movements of the eyebrows and changes in the mouth corners. As output, these analysis results are quantified as concentration and understanding levels and stored in the database.

[0292] Step 3:

[0293] The server performs speech analysis using audio data. It receives recorded audio as input and extracts speech features by applying a speech recognition algorithm. Specifically, it analyzes changes in the instructor's speaking tempo and volume and compares them with the learner's reactions. As output, an evaluation of the effectiveness of the instruction content is quantified and integrated with the results of facial expression analysis.

[0294] Step 4:

[0295] The device generates real-time notifications based on the analysis results. It receives analysis data as input and analyzes and makes decisions based on set thresholds. For example, if the "concentration level" falls below a predetermined value, it notifies the instructor that "Person A's concentration is declining." Actionable feedback is displayed on the device screen as output.

[0296] Step 5:

[0297] The server automatically generates evaluation reports based on the integrated data after each lesson. It receives the aggregated analytical data as input and processes it to visualize the progress of each learner's understanding and concentration levels. Specifically, the report outputs areas where understanding has improved and areas requiring improvement. The output is provided in a format viewable by instructors, which can be used to improve future lessons.

[0298] Step 6:

[0299] The server generates specific suggestions for improving teaching methods based on recorded data. Using analysis results stored in the database as input, it employs an AI model to derive optimal improvement suggestions. The output provides instructors with suggestions such as "increase the use of visual aids," which can be used to plan future lessons.

[0300] (Application Example 1)

[0301] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0302] In educational environments and families, it is required to dynamically grasp the understanding and concentration levels of learners and provide real-time feedback. However, with conventional methods, there are problems such as delays in feedback in educational settings and a tendency to lack data for providing appropriate guidance to individual learners. Also, in families, there is a current situation where it is difficult for guardians to accurately grasp the learning status of children and take appropriate actions.

[0303] The specific processing by the specific processing unit 290 of the data processing apparatus 12 in Application Example 1 is realized by the following respective means.

[0304] In this invention, the server includes means for acquiring information for analyzing the understanding and concentration levels of learners in an educational environment, analysis means for evaluating the emotional state of learners and the effect of the educator's explanation to the learners using the information acquired by the means, and means for notifying the educator in real time based on the evaluation obtained by the analysis means. Thereby, it becomes possible to quickly and accurately grasp the state of learners, and for educators and guardians to take immediate actions.

[0305] "Educational environment" is a general term for the places and situations where educational activities for learners are carried out.

[0306] "Learner" refers to a person who receives education in order to acquire knowledge and skills.

[0307] "Understanding level" is a measure indicating the degree to which a learner understands the learning content.

[0308] "Concentration level" is a measure representing the degree to which a learner pays attention to learning activities.

[0309] "Means for acquiring information" refers to devices or methods for collecting data related to learners.

[0310] "Analysis means" refers to the processes and techniques used to analyze acquired information and evaluate the learner's state.

[0311] "Means of notification" refers to methods or devices for conveying specific information to educators or parents based on analysis results.

[0312] "Guardian" refers to an adult who is responsible for the upbringing of a learner.

[0313] "Means of generating suggestions" refers to methods and processes for creating suggestions that help improve teaching methods based on acquired and analyzed data.

[0314] This system aims to analyze learners' comprehension and concentration levels in educational and home environments and provide real-time feedback to educators and parents. The following describes its configuration and operating procedures in detail.

[0315] The server uses cameras and microphones installed in educational and home environments to acquire learners' facial expressions, posture, and voice data. This allows for the collection of data tailored to the situation in real time. The hardware uses a standard webcam as the camera and a directional microphone as the microphone.

[0316] The collected data is processed on the server. For image data, facial recognition is performed using OpenCV and TensorFlow. Audio data is analyzed using PyDub to evaluate the intonation and tempo of the speech. This analysis allows us to determine whether the learner understands or is confused.

[0317] Based on the analysis results, the device sends notifications to educators and parents. Using the Pushbullet API, messages are sent in real time to smartphones and tablets. For example, a notification might appear stating, "Learner A may be having difficulty understanding the material." This information can be used to immediately adjust teaching methods.

[0318] Furthermore, after a learning session ends, the server automatically generates a detailed report based on the collected and analyzed data. This report is provided to educators as a reference for improving their teaching methods and can also be used to support learning at home.

[0319] As a concrete example, when a student is working on a math assignment at home, facial recognition can detect signs of confusion, such as wrinkles between the eyebrows or changes in eye movement. The parent is then immediately notified that "Student A appears to be having difficulty understanding the material."

[0320] An example of a prompt to input into the generating AI model is: "Analyze my child's facial expressions while they are studying math, determine if they are having difficulty understanding, and send a notification."

[0321] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0322] Step 1:

[0323] The server collects learners' facial expressions, posture, and voice data in real time through installed cameras and microphones. This input data includes changes in facial expressions and vocal intonation, and is used as foundational data to understand the learners' emotional state. Specifically, the camera captures facial feature points, and the microphone records the tone and volume of the voice.

[0324] Step 2:

[0325] The server performs facial recognition processing on the acquired image data using OpenCV and TensorFlow. This process extracts feature points from the learner's face from the input data and analyzes what emotions the input facial image indicates. For example, if there are wrinkles between the eyebrows, it will be determined that the person is confused. The output will be the analysis result of the learner's emotional state.

[0326] Step 3:

[0327] The server uses PyDub to perform speech analysis on the audio data. This analysis examines the intonation and tempo of the input audio data to determine the learner's level of concentration and comprehension. Specifically, if the tone of voice suddenly becomes flat, it is determined that the learner's concentration may have been interrupted. The output of this analysis represents the results of the comprehension and concentration assessment.

[0328] Step 4:

[0329] Based on the analysis results, the device uses the Pushbullet API to send notifications to educators and parents. In this step, the device displays the analyzed results of the learner's comprehension and concentration levels as a message. For example, a notification saying "Learner A may be having difficulty understanding" is sent to the parent's smartphone. The input is the analysis results, and the output is the notification message.

[0330] Step 5:

[0331] The server automatically generates a comprehensive evaluation report based on the data collected after each learning session. This report includes information on when the learner understood the material or when their concentration wavered, providing valuable insights for improving future teaching methods. Specifically, the server organizes the collected data chronologically and summarizes the evaluation results. The input is the previously analyzed data, and the output is the evaluation report.

[0332] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0333] This invention combines an emotion engine with a system for analyzing learners' comprehension and concentration levels in an educational environment. This system consists of an information acquisition device, an analysis device, a notification device, a recording device, a suggestion generation device, and an emotion engine.

[0334] The process of information acquisition and emotion recognition

[0335] The server collects learners' facial expressions and voice data in real time through cameras and microphones installed in the classroom. The emotion engine uses this data to recognize the learners' emotional state and identify emotions such as "interesting" or "confusing."

[0336] Data analysis and coaching feedback process

[0337] The server inputs data, including emotional states recognized by the emotion engine, into the analysis device. The analysis device evaluates the relationship between the learner's emotional state, comprehension level, and concentration level, and makes a comprehensive judgment on the effectiveness of the instruction. As a result, it is possible to understand how the learner is receiving the lesson content. For example, if a learner is "confused," it is determined that the teaching method at that point should be reconsidered.

[0338] Real-time notifications and the process of improving instruction

[0339] The device provides real-time feedback to instructors based on the analysis results. For example, a message such as "Ms. E seems a little confused" might appear on the instructor's device, allowing the instructor to adjust their teaching method immediately. Feedback is also provided if certain emotions are repeatedly observed.

[0340] Recording and proposal generation process

[0341] The server accumulates past emotional data and analysis results to evaluate changes in learners' states over the long term. Based on this, it generates suggestions for optimizing teaching methods. For example, if it is confirmed that some learners' comprehension improves with audiovisual stimuli, the use of visual materials will be suggested.

[0342] This system enables the improvement of educational quality by conducting multifaceted analysis, including learners' emotions. Instructors can grasp learners' emotions in real time and gain concrete means to implement appropriate instruction.

[0343] The following describes the processing flow.

[0344] Step 1:

[0345] The server uses cameras and microphones in the classroom to collect learners' facial expressions and audio data in real time. The cameras capture each learner's facial expressions, and the microphones record the tone and intonation of their individual voices. This data is temporarily stored in a database in its raw state.

[0346] Step 2:

[0347] The server preprocesses the collected data. Image data undergoes image filtering to identify facial features and extract expressions. Audio data is processed through noise filtering and audio frame extraction, and then formatted to be suitable for audio analysis methods.

[0348] Step 3:

[0349] The server supplies pre-processed data to the emotion engine. The emotion engine uses machine learning algorithms to analyze emotions from facial expressions and voice, determining emotional states such as "joy," "confusion," and "distraction." This enables instantaneous diagnosis of the learner's emotions.

[0350] Step 4:

[0351] The device receives sentiment analysis results sent from the server and notifies the instructor. Specifically, it displays a message such as, "Student F is confused by the lesson content," prompting the instructor to adjust the lesson's pace and explanation methods. This notification enables flexible responses during the lesson.

[0352] Step 5:

[0353] The server tracks and stores emotional data and associated analytical information over long periods. The stored data is organized in a format that allows for the evaluation of emotional changes over time. As a result, instructors can understand the temporal evolution of learners' emotions and use this information to improve teaching methods.

[0354] Step 6:

[0355] The server generates specific suggestions for optimizing teaching methods based on the combined data. These suggestions include actionable improvements, such as "Increasing visual aids can improve comprehension," providing practical guidelines for instructors.

[0356] (Example 2)

[0357] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0358] In educational settings, it is difficult to instantly grasp learners' levels of understanding and concentration and provide appropriate instruction; therefore, efficient methods are needed to maximize learning effectiveness. In particular, there is a need to develop a system that allows instructors to respond quickly by providing real-time feedback on learners' emotional states.

[0359] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0360] In this invention, the server includes means for collecting video and audio from the environment to identify the learner's emotions; means for analyzing the collected video and audio and evaluating the learner's emotions and learning effectiveness using a specified algorithm; and means for providing real-time feedback to the instructor according to the evaluated emotional state. This makes it possible to instantly grasp the learner's level of understanding and emotional state and improve the quality of instruction.

[0361] A "learner" is someone who seeks to acquire knowledge and skills in an educational environment.

[0362] "Emotions" refer to the mental state of a learner and include states such as interest, confusion, happiness, and anxiety.

[0363] "Video" refers to dynamic or static image data that visually represents the learner's facial expressions and actions.

[0364] "Speech" refers to acoustic signals, including the learner's utterances and tone of voice.

[0365] An "algorithm" is a set of procedures or computational steps for solving a specific problem, and it is used as a means of analyzing learners' emotions.

[0366] "Feedback" is information provided to instructors regarding the learner's current state, which allows them to adjust their teaching methods accordingly.

[0367] A "threshold" is a specific numerical value or level set as a standard for evaluating a learner's level of understanding and concentration.

[0368] A "suggestion" is information that provides instructors with specific guidance on improving and optimizing teaching methods, based on recorded data and analysis results.

[0369] This invention is a system for analyzing learners' comprehension and concentration levels in an educational environment and optimizing teaching methods. The system consists of three main components: a server, terminals, and users.

[0370] First, the server uses cameras and microphones placed in the classroom to collect learners' facial expressions and voice data in real time. This hardware operates continuously to facilitate data acquisition. The collected data is processed by an emotion engine to analyze the learners' emotional state. This emotion engine is equipped with emotion recognition algorithms and has advanced capabilities to analyze subtle changes in facial expressions and tone of voice.

[0371] Next, based on this analyzed data, the device provides real-time feedback to the instructor. The device immediately notifies the instructor according to the analysis results, displaying specific messages such as, "Student E is confused." This allows the instructor to immediately adopt teaching methods that respond to the learner's situation.

[0372] Furthermore, the server accumulates this data and analysis results over the long term, continuously evaluating changes in learners' situations. Based on this long-term data, the server can generate specific suggestions for improving teaching methods. This allows instructors to obtain valuable information for continuously optimizing their teaching methods.

[0373] For example, if data reveals that visual learning materials improve the comprehension of certain learners, then teaching methods can be proposed. Instructors can then use these suggestions to further improve the quality of their instruction.

[0374] An example of a prompt for a generative AI model is: "Please explain how to analyze learners' emotions in a classroom and evaluate their comprehension and concentration levels. Also, show how this data can be used to optimize instruction."

[0375] This system enables data-driven instruction to improve the quality of education.

[0376] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0377] Step 1:

[0378] The server uses cameras and microphones installed in the classroom to collect learners' facial expressions and voice data as input. Specifically, the cameras capture video at several frames per second and send it to the server as digital data. The microphones record learners' speech and ambient sounds in real time and send them to the server in digital format. This generates a petabyte-scale dataset that reflects the learners' state.

[0379] Step 2:

[0380] The server takes the collected facial expression and audio data as input and entrusts the processing to the emotion engine. Specifically, the emotion engine uses a facial recognition algorithm to analyze the facial features of each frame and uses speech recognition technology to evaluate the tone of voice and emotion. As output, the server generates emotion states labeled as "interesting" or "confused." These labels indicate the learner's state in real time.

[0381] Step 3:

[0382] The server sends emotional state data obtained from the emotion engine as input to the analysis device. The analysis device uses data mining techniques to compare this data with past data and evaluate the learner's level of understanding and concentration. As output, evaluation results regarding the effectiveness of specific teaching methods are generated. This prepares the server to receive feedback on the effectiveness of the instruction as numerical values ​​and evaluation comments.

[0383] Step 4:

[0384] The terminal receives evaluation results from the server as input and provides real-time feedback to the instructor. Specifically, the terminal displays a warning message such as "Person E is confused," allowing the instructor to immediately adjust their teaching methods based on the feedback. The output is information provided to the instructor and improvements to the teaching based on that information. This rapid feedback enables instructors to respond immediately to provide appropriate instruction.

[0385] Step 5:

[0386] The server records all data and analysis results over the long term as input. Specifically, it stores data in a database and organizes it to be useful for subsequent analysis. Based on this data, a suggestion generator generates suggestions for optimizing teaching methods. As output, instructors are provided with concrete suggestions for long-term educational improvement. This establishes a sustainable improvement cycle for improving the quality of education.

[0387] (Application Example 2)

[0388] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0389] In educational settings, it is difficult to appropriately assess learners' comprehension and concentration levels and provide real-time feedback. Traditional systems fail to adequately grasp learners' emotional states, making it difficult to improve teaching methods to suit individual learners. Furthermore, while there is a need to respond quickly to situations where learners are confused and provide appropriate educational materials, there is a lack of effective means to achieve this.

[0390] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0391] In this invention, the server includes means for acquiring information to analyze learners' awareness and concentration levels in an educational environment, means for evaluating learners' emotional states and the effectiveness of instructors' explanations using the acquired information, and means for notifying instructors in real time based on the evaluation and providing interactive educational materials to promote understanding among learners. This makes it possible to instantly grasp learners' emotional states and provide optimal teaching methods and materials for each individual learner.

[0392] The term "educational environment" refers to the place or situation in which learners acquire knowledge and skills, and includes physical classrooms and digital platforms.

[0393] The term "learner" refers to individuals who seek to acquire knowledge and skills within an educational environment.

[0394] "Recognition level" is an indicator that shows how well learners understand the educational content.

[0395] "Concentration" refers to the state in which learners are paying attention to educational activities.

[0396] "Emotional state" refers to the psychological state or reaction that learners exhibit during learning, and includes specific emotions such as "interesting" or "confusing."

[0397] "Interactive educational materials" refer to learning content and materials that allow learners to actively participate and interact with each other.

[0398] "Evaluation" refers to the process of analyzing learners' comprehension, concentration levels, and emotional states based on acquired data, and measuring the effectiveness of teaching methods.

[0399] "Notification" refers to the act of communicating information to instructors based on analysis results.

[0400] "Feedback" refers to information provided to improve teaching methods and enhance educational effectiveness, based on evaluation results, learner responses, and other factors.

[0401] The system implementing this invention analyzes learners' emotional states and comprehension levels in real time within an educational environment and provides appropriate feedback. The server uses cameras and microphones installed in the educational space to collect learners' facial expressions and voices in real time. This collected data is analyzed through emotion recognition software, such as the Microsoft Azure Emotion API, to detect the learners' emotional states.

[0402] The device receives analysis results from the server and notifies the instructor in real time with feedback tailored to the learner's emotional state. This notification includes specific advice for the instructor to adjust their teaching methods according to the learner's condition.

[0403] Furthermore, the server records changes in learners' comprehension and emotional states, and generates suggestions for evaluating long-term learning effectiveness. These suggestions include guidelines on how interactive educational materials should be used.

[0404] For example, in a physics lesson in a classroom, the server might detect that a student is experiencing "confusion" when faced with a difficult problem. The terminal would immediately notify the instructor with a message such as, "Some students are confused. Please try explaining using concrete examples." This notification allows the instructor to flexibly adjust their teaching methods.

[0405] Possible prompt statements for input to a generative AI model include the following:

[0406] "How can you engage students who aren't interested in historical discussions?"

[0407] "When we detect that students are confused while solving math problems, what kind of interactive learning materials can we provide?"

[0408] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0409] Step 1:

[0410] The server collects learners' facial expressions and audio data in real time through cameras and microphones installed in the educational space. Inputs include video streams and audio streams, which are stored in a database. The operations performed at this stage involve video capture by the cameras and audio recording by the microphones.

[0411] Step 2:

[0412] The server sends the collected facial expression and voice data to emotion recognition software (e.g., Microsoft Azure Emotion API) to analyze the learner's emotional state. The input is in the form of raw data, which the emotion recognition software processes to generate outputs representing the learner's intuitive psychological state, such as labels like "interesting" or "confused." Data preprocessing and emotion analysis are the main operations in this step.

[0413] Step 3:

[0414] The server inputs emotional state data obtained from emotion recognition software into an analysis device to evaluate learners' comprehension and concentration levels. At this stage, data calculations are performed to correlate emotional states with the progress of the lesson and generate evaluation results. The output provides individual learners' evaluations of comprehension and concentration levels.

[0415] Step 4:

[0416] The terminal notifies instructors of evaluation results in real time, helping them select appropriate teaching methods. The input is the evaluation results sent from the server, and the output is a specific feedback message presented to the instructor. The key feature here is that the notification system operates in real time.

[0417] Step 5:

[0418] The server stores evaluation results and emotional state records in a database and generates suggestions for improving teaching methods based on that data. Long-term learner data history is used as input, and actionable suggestions are generated as output. The main operations in this step are data analysis and suggestion generation.

[0419] Step 6:

[0420] The terminal provides the instructor with generated suggestions and directly offers learners interactive educational materials tailored to specific situations as needed. Input is data from the suggestion generator, and output is specific educational material. The operation here is an appropriate response based on the learner's situation.

[0421] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0422] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0423] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0424] [Third Embodiment]

[0425] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0426] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0427] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0428] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0429] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0430] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0431] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0432] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0433] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0434] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0435] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0436] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0437] This invention relates to a system that analyzes learners' comprehension and concentration levels in an educational environment and provides real-time feedback to instructors. This system consists of an information acquisition device, an analysis device, a notification device, a recording device, and a suggestion generation device.

[0438] Information acquisition process

[0439] The server collects learners' facial expressions, posture, and voice data via cameras and microphones installed in the classroom. This provides foundational data to understand the learners' current emotional state. For example, the camera detects facial movements when a learner frowns and records this as a change in facial expression.

[0440] Data analysis process

[0441] The server processes the collected data using advanced image and audio analysis algorithms. Image data is analyzed using facial recognition technology to determine whether the learner understands or is confused. Audio data is analyzed using audio analysis technology to extract characteristics of the instructor's speaking style (e.g., intonation and tempo) and evaluate the effectiveness of the instruction.

[0442] Real-time notification process

[0443] The device notifies the instructor in real time if it detects a decrease in the learner's concentration or that they are having difficulty understanding the material. For example, a message such as "Student D may be having difficulty understanding" might appear on the instructor's device. This information can then be used by the instructor to adjust their teaching methods on the spot.

[0444] Report generation process

[0445] After the lesson ends, the server automatically generates a comprehensive evaluation report based on the data collected. The report specifically shows which learners made progress in understanding and at what point their concentration wavered. This allows instructors to review the entire lesson and use the results to improve their teaching later.

[0446] Proposal generation process

[0447] The server is a device that uses recorded data and analysis results to generate improvement suggestions for instructors. This allows instructors to obtain specific, actionable directions for further deepening their teaching methods. For example, it might provide suggestions such as, "Incorporating visual aids into lessons may improve comprehension."

[0448] This system enables the integrated and dynamic assessment and response to learners' understanding and concentration within the educational environment, thereby promoting improvements in the quality of education.

[0449] The following describes the processing flow.

[0450] Step 1:

[0451] The server collects video and audio data in real time through cameras and microphones installed in the classroom. The cameras capture each student's face and expressions, and the microphones record the instructor's voice. During this process, the image data is processed frame by frame, and the audio data is converted to the appropriate format.

[0452] Step 2:

[0453] The server preprocesses the collected image and audio data. This preprocessing involves performing face recognition and expression extraction using the image data, and removing noise from the audio data to generate clear audio clips. This process prepares the data for analysis.

[0454] Step 3:

[0455] The server inputs preprocessed data into an AI model and performs analysis using a deep learning algorithm. The analysis classifies the learner's emotional state from their facial expressions and evaluates their concentration and comprehension levels. It also analyzes the instructor's speaking speed and intonation from audio data to determine if the instruction is effective.

[0456] Step 4:

[0457] The device receives analysis results from the server and sends real-time notifications to the instructor. These notifications include alerts if the learner's comprehension is declining or if their speaking style needs improvement. These notifications appear as pop-ups on the instructor's device.

[0458] Step 5:

[0459] At the end of each lesson, the server generates a report based on an overall assessment of comprehension and concentration levels derived from analysis. This report includes each learner's responses and changes in concentration at specific points during the lesson. Instructors can use this information to improve future lessons.

[0460] Step 6:

[0461] The server generates specific suggestions for improving teaching methods based on accumulated data. These suggestions include concrete actions to improve learner responses based on past data, supporting the growth of instructors.

[0462] (Example 1)

[0463] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0464] In educational settings, there is a lack of means to accurately grasp learners' levels of understanding and concentration in real time and to quickly convey that information to instructors. Furthermore, there is a need for effective methods to appropriately analyze changes in learners' emotional states and learning progress and use that information to improve instruction. This presents a challenge in that it is difficult for instructors to adjust their teaching methods to be optimal for each learner in a timely manner.

[0465] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0466] In this invention, the server includes means for acquiring the learner's emotional state, posture, and voice information in the learning environment; means for analyzing the learner's emotional state and the characteristics of the instructor's explanation using the image and voice information acquired by the means; and means for providing notifications that enable real-time adjustment of the teaching method based on the analysis results. This makes it possible to grasp the learner's level of understanding and concentration in real time and dynamically adjust the teaching method based on that.

[0467] The term "learning environment" refers to the place or space where education takes place, as well as the equipment and technology used within it.

[0468] A "learner" refers to an individual who participates in educational activities with the aim of acquiring knowledge and skills.

[0469] "Emotional state" refers to the psychological expressions and emotional states exhibited by learners, which can be observed through changes in facial expressions and voice.

[0470] "Postural information" refers to data related to the learner's body position and movement, and is used to determine their level of concentration and interest.

[0471] "Auditory information" refers to sound data collected within the learning environment, encompassing the speech of learners and instructors, as well as ambient sounds.

[0472] "Analysis results" refer to the conclusions and insights derived from processing the collected data.

[0473] "Notifications" refer to information and messages generated based on analysis and provided to instructors.

[0474] "Teaching methods" refer to the techniques and processes by which an instructor provides education to learners.

[0475] "Recording" refers to the act of saving data and information so that it can be referenced or analyzed later, or the result of doing so.

[0476] "Improvement suggestions" refer to specific proposals provided to improve the quality of education, based on recorded data and analysis.

[0477] This invention aims to implement a system that analyzes learners' comprehension and concentration levels in real time within educational settings and provides feedback to instructors. The details are described below.

[0478] The server utilizes cameras and microphones installed within the learning environment to collect learners' facial expressions, posture, and voice information. This data forms the basis for understanding the learners' emotional state at the time of acquisition.

[0479] Next, the server processes the collected data using advanced image and audio analysis software. Image data is analyzed based on facial expression recognition technology to identify how well the learner understands or is confused. Meanwhile, audio data is analyzed for changes in tone and tempo to evaluate the characteristics of the instructor's explanation.

[0480] The terminal notifies the instructor in real time based on the judgment results obtained from the server. This notification is displayed to the instructor in the form of, for example, "Student D may be having difficulty understanding." The instructor uses this information to adjust the lesson plan and approach on the spot.

[0481] Furthermore, after the lesson ends, the server automatically generates an evaluation report based on the analysis results. This report details which learners showed difficulties with understanding or concentration and at what point, and is used as feedback for instructors.

[0482] Furthermore, based on the recorded data and evaluation results, suggestions for improving teaching methods are generated. For example, a suggestion might be made such as, "Incorporating visual materials into lessons may improve learners' comprehension."

[0483] This system will promote increased efficiency and improved quality of instruction in educational settings.

[0484] Examples of prompt statements include the following:

[0485] "Could you provide an example of a notification message sent to an instructor if they determine that a student is losing focus during class?"

[0486] "Please provide three suggestions for improving learners' comprehension."

[0487] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0488] Step 1:

[0489] The server acquires the learner's facial expressions, posture, and voice information via the camera and microphone within the learning environment. It receives camera video and audio data as input and uses this to aggregate basic data about the learner's state in real time. Specifically, the image sensor captures facial expression data, and the voice sensor records the learner's speech and ambient sounds. The collected raw data is sent to the analysis module as output.

[0490] Step 2:

[0491] The server performs facial recognition using collected image data. It receives raw data as input and applies image processing algorithms to classify the learner's emotional state. For example, it determines states such as "understanding," "confused," and "indifferent" from eyebrow movements and changes in the mouth. As output, these analysis results are quantified as levels of concentration and comprehension and stored in a database.

[0492] Step 3:

[0493] The server performs speech analysis using audio data. It receives recorded audio as input and extracts speech features by applying a speech recognition algorithm. Specifically, it analyzes changes in the instructor's speaking tempo and volume and compares them with the learner's reactions. As output, an evaluation of the effectiveness of the instruction content is quantified and integrated with the results of facial expression analysis.

[0494] Step 4:

[0495] The device generates real-time notifications based on the analysis results. It receives analysis data as input and analyzes and makes decisions based on set thresholds. For example, if the "concentration level" falls below a predetermined value, it notifies the instructor that "Person A's concentration is declining." Actionable feedback is displayed on the device screen as output.

[0496] Step 5:

[0497] The server automatically generates evaluation reports based on the integrated data after each lesson. It receives the aggregated analytical data as input and processes it to visualize the progress of each learner's understanding and concentration levels. Specifically, the report outputs areas where understanding has improved and areas requiring improvement. The output is provided in a format viewable by instructors, which can be used to improve future lessons.

[0498] Step 6:

[0499] The server generates specific suggestions for improving teaching methods based on recorded data. Using analysis results stored in the database as input, it employs an AI model to derive optimal improvement suggestions. The output provides instructors with suggestions such as "increase the use of visual aids," which can be used to plan future lessons.

[0500] (Application Example 1)

[0501] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0502] In educational settings and at home, there is a growing need to dynamically grasp learners' comprehension and concentration levels and provide real-time feedback. However, traditional methods have faced challenges such as delayed feedback in educational settings and a lack of data necessary for providing appropriate guidance to individual learners. Furthermore, at home, it is currently difficult for parents to accurately grasp their children's learning progress and take appropriate action.

[0503] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0504] In this invention, the server includes means for acquiring information for analyzing the learner's level of understanding and concentration in an educational environment; analysis means for evaluating the learner's emotional state and the effectiveness of the educator's explanations to the learner using the information acquired by the means; and means for notifying the educator in real time based on the evaluation obtained by the analysis means. This enables the learner's condition to be grasped quickly and accurately, allowing educators and parents to take immediate action.

[0505] "Educational environment" is a general term for the places and circumstances in which educational activities are conducted for learners.

[0506] A "learner" refers to someone who receives education in order to acquire knowledge and skills.

[0507] "Comprehension level" is a measure that indicates the extent to which a learner understands the learning material.

[0508] "Concentration level" is a measure that indicates how much attention a learner is paying to a learning activity.

[0509] "Means of acquiring information" refers to devices and methods for collecting data about learners.

[0510] "Analysis means" refers to the processes and techniques used to analyze acquired information and evaluate the learner's state.

[0511] "Means of notification" refers to methods or devices for conveying specific information to educators or parents based on analysis results.

[0512] "Guardian" refers to an adult who is responsible for the upbringing of a learner.

[0513] "Means of generating suggestions" refers to methods and processes for creating suggestions that help improve teaching methods based on acquired and analyzed data.

[0514] This system aims to analyze learners' comprehension and concentration levels in educational and home environments and provide real-time feedback to educators and parents. The following describes its configuration and operating procedures in detail.

[0515] The server uses cameras and microphones installed in educational and home environments to acquire learners' facial expressions, posture, and voice data. This allows for the collection of data tailored to the situation in real time. The hardware uses a standard webcam as the camera and a directional microphone as the microphone.

[0516] The collected data is processed on the server. For image data, facial recognition is performed using OpenCV and TensorFlow. Audio data is analyzed using PyDub to evaluate the intonation and tempo of the speech. This analysis allows us to determine whether the learner understands or is confused.

[0517] Based on the analysis results, the device sends notifications to educators and parents. Using the Pushbullet API, messages are sent in real time to smartphones and tablets. For example, a notification might appear stating, "Learner A may be having difficulty understanding the material." This information can be used to immediately adjust teaching methods.

[0518] Furthermore, after a learning session ends, the server automatically generates a detailed report based on the collected and analyzed data. This report is provided to educators as a reference for improving their teaching methods and can also be used to support learning at home.

[0519] As a concrete example, when a student is working on a math assignment at home, facial recognition can detect signs of confusion, such as wrinkles between the eyebrows or changes in eye movement. The parent is then immediately notified that "Student A appears to be having difficulty understanding the material."

[0520] An example of a prompt to input into the generating AI model is: "Analyze my child's facial expressions while they are studying math, determine if they are having difficulty understanding, and send a notification."

[0521] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0522] Step 1:

[0523] The server collects learners' facial expressions, posture, and voice data in real time through installed cameras and microphones. This input data includes changes in facial expressions and vocal intonation, and is used as foundational data to understand the learners' emotional state. Specifically, the camera captures facial feature points, and the microphone records the tone and volume of the voice.

[0524] Step 2:

[0525] The server performs facial recognition processing on the acquired image data using OpenCV and TensorFlow. This process extracts feature points from the learner's face from the input data and analyzes what emotions the input facial image indicates. For example, if there are wrinkles between the eyebrows, it will be determined that the person is confused. The output will be the analysis result of the learner's emotional state.

[0526] Step 3:

[0527] The server uses PyDub to perform speech analysis on the audio data. This analysis examines the intonation and tempo of the input audio data to determine the learner's level of concentration and comprehension. Specifically, if the tone of voice suddenly becomes flat, it is determined that the learner's concentration may have been interrupted. The output of this analysis represents the results of the comprehension and concentration assessment.

[0528] Step 4:

[0529] Based on the analysis results, the device uses the Pushbullet API to send notifications to educators and parents. In this step, the device displays the analyzed results of the learner's comprehension and concentration levels as a message. For example, a notification saying "Learner A may be having difficulty understanding" is sent to the parent's smartphone. The input is the analysis results, and the output is the notification message.

[0530] Step 5:

[0531] The server automatically generates a comprehensive evaluation report based on the data collected after each learning session. This report includes information on when the learner understood the material or when their concentration wavered, providing valuable insights for improving future teaching methods. Specifically, the server organizes the collected data chronologically and summarizes the evaluation results. The input is the previously analyzed data, and the output is the evaluation report.

[0532] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0533] This invention combines an emotion engine with a system for analyzing learners' comprehension and concentration levels in an educational environment. This system consists of an information acquisition device, an analysis device, a notification device, a recording device, a suggestion generation device, and an emotion engine.

[0534] The process of information acquisition and emotion recognition

[0535] The server collects learners' facial expressions and voice data in real time through cameras and microphones installed in the classroom. The emotion engine uses this data to recognize the learners' emotional state and identify emotions such as "interesting" or "confusing."

[0536] Data analysis and coaching feedback process

[0537] The server inputs data, including emotional states recognized by the emotion engine, into the analysis device. The analysis device evaluates the relationship between the learner's emotional state, comprehension level, and concentration level, and makes a comprehensive judgment on the effectiveness of the instruction. As a result, it is possible to understand how the learner is receiving the lesson content. For example, if a learner is "confused," it is determined that the teaching method at that point should be reconsidered.

[0538] Real-time notifications and the process of improving instruction

[0539] The device provides real-time feedback to instructors based on the analysis results. For example, a message such as "Ms. E seems a little confused" might appear on the instructor's device, allowing the instructor to adjust their teaching method immediately. Feedback is also provided if certain emotions are repeatedly observed.

[0540] Recording and proposal generation process

[0541] The server accumulates past emotional data and analysis results to evaluate changes in learners' states over the long term. Based on this, it generates suggestions for optimizing teaching methods. For example, if it is confirmed that some learners' comprehension improves with audiovisual stimuli, the use of visual materials will be suggested.

[0542] This system enables the improvement of educational quality by conducting multifaceted analysis, including learners' emotions. Instructors can grasp learners' emotions in real time and gain concrete means to implement appropriate instruction.

[0543] The following describes the processing flow.

[0544] Step 1:

[0545] The server uses cameras and microphones in the classroom to collect learners' facial expressions and audio data in real time. The cameras capture each learner's facial expressions, and the microphones record the tone and intonation of their individual voices. This data is temporarily stored in a database in its raw state.

[0546] Step 2:

[0547] The server preprocesses the collected data. Image data undergoes image filtering to identify facial features and extract expressions. Audio data is processed through noise filtering and audio frame extraction, and then formatted to be suitable for audio analysis methods.

[0548] Step 3:

[0549] The server supplies pre-processed data to the emotion engine. The emotion engine uses machine learning algorithms to analyze emotions from facial expressions and voice, determining emotional states such as "joy," "confusion," and "distraction." This enables instantaneous diagnosis of the learner's emotions.

[0550] Step 4:

[0551] The device receives sentiment analysis results sent from the server and notifies the instructor. Specifically, it displays a message such as, "Student F is confused by the lesson content," prompting the instructor to adjust the lesson's pace and explanation methods. This notification enables flexible responses during the lesson.

[0552] Step 5:

[0553] The server tracks and stores emotional data and associated analytical information over long periods. The stored data is organized in a format that allows for the evaluation of emotional changes over time. As a result, instructors can understand the temporal evolution of learners' emotions and use this information to improve teaching methods.

[0554] Step 6:

[0555] The server generates specific suggestions for optimizing teaching methods based on the combined data. These suggestions include actionable improvements, such as "Increasing visual aids can improve comprehension," providing practical guidelines for instructors.

[0556] (Example 2)

[0557] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0558] In educational settings, it is difficult to instantly grasp learners' levels of understanding and concentration and provide appropriate instruction; therefore, efficient methods are needed to maximize learning effectiveness. In particular, there is a need to develop a system that allows instructors to respond quickly by providing real-time feedback on learners' emotional states.

[0559] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0560] In this invention, the server includes means for collecting video and audio from the environment to identify the learner's emotions; means for analyzing the collected video and audio and evaluating the learner's emotions and learning effectiveness using a specified algorithm; and means for providing real-time feedback to the instructor according to the evaluated emotional state. This makes it possible to instantly grasp the learner's level of understanding and emotional state and improve the quality of instruction.

[0561] A "learner" is someone who seeks to acquire knowledge and skills in an educational environment.

[0562] "Emotions" refer to the mental state of a learner and include states such as interest, confusion, happiness, and anxiety.

[0563] "Video" refers to dynamic or static image data that visually represents the learner's facial expressions and actions.

[0564] "Speech" refers to acoustic signals, including the learner's utterances and tone of voice.

[0565] An "algorithm" is a set of procedures or computational steps for solving a specific problem, and it is used as a means of analyzing learners' emotions.

[0566] "Feedback" is information provided to instructors regarding the learner's current state, which allows them to adjust their teaching methods accordingly.

[0567] A "threshold" is a specific numerical value or level set as a standard for evaluating a learner's level of understanding and concentration.

[0568] A "suggestion" is information that provides instructors with specific guidance on improving and optimizing teaching methods, based on recorded data and analysis results.

[0569] This invention is a system for analyzing learners' comprehension and concentration levels in an educational environment and optimizing teaching methods. The system consists of three main components: a server, terminals, and users.

[0570] First, the server uses cameras and microphones placed in the classroom to collect learners' facial expressions and voice data in real time. This hardware operates continuously to facilitate data acquisition. The collected data is processed by an emotion engine to analyze the learners' emotional state. This emotion engine is equipped with emotion recognition algorithms and has advanced capabilities to analyze subtle changes in facial expressions and tone of voice.

[0571] Next, based on this analyzed data, the device provides real-time feedback to the instructor. The device immediately notifies the instructor according to the analysis results, displaying specific messages such as, "Student E is confused." This allows the instructor to immediately adopt teaching methods that respond to the learner's situation.

[0572] Furthermore, the server accumulates this data and analysis results over the long term, continuously evaluating changes in learners' situations. Based on this long-term data, the server can generate specific suggestions for improving teaching methods. This allows instructors to obtain valuable information for continuously optimizing their teaching methods.

[0573] For example, if data reveals that visual learning materials improve the comprehension of certain learners, then teaching methods can be proposed. Instructors can then use these suggestions to further improve the quality of their instruction.

[0574] An example of a prompt for a generative AI model is: "Please explain how to analyze learners' emotions in a classroom and evaluate their comprehension and concentration levels. Also, show how this data can be used to optimize instruction."

[0575] This system enables data-driven instruction to improve the quality of education.

[0576] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0577] Step 1:

[0578] The server uses cameras and microphones installed in the classroom to collect learners' facial expressions and voice data as input. Specifically, the cameras capture video at several frames per second and send it to the server as digital data. The microphones record learners' speech and ambient sounds in real time and send them to the server in digital format. This generates a petabyte-scale dataset that reflects the learners' state.

[0579] Step 2:

[0580] The server takes the collected facial expression and audio data as input and entrusts the processing to the emotion engine. Specifically, the emotion engine uses a facial recognition algorithm to analyze the facial features of each frame and uses speech recognition technology to evaluate the tone of voice and emotion. As output, the server generates emotion states labeled as "interesting" or "confused." These labels indicate the learner's state in real time.

[0581] Step 3:

[0582] The server sends emotional state data obtained from the emotion engine as input to the analysis device. The analysis device uses data mining techniques to compare this data with past data and evaluate the learner's level of understanding and concentration. As output, evaluation results regarding the effectiveness of specific teaching methods are generated. This prepares the server to receive feedback on the effectiveness of the instruction as numerical values ​​and evaluation comments.

[0583] Step 4:

[0584] The terminal receives evaluation results from the server as input and provides real-time feedback to the instructor. Specifically, the terminal displays a warning message such as "Person E is confused," allowing the instructor to immediately adjust their teaching methods based on the feedback. The output is information provided to the instructor and improvements to the teaching based on that information. This rapid feedback enables instructors to respond immediately to provide appropriate instruction.

[0585] Step 5:

[0586] The server records all data and analysis results over the long term as input. Specifically, it stores data in a database and organizes it to be useful for subsequent analysis. Based on this data, a suggestion generator generates suggestions for optimizing teaching methods. As output, instructors are provided with concrete suggestions for long-term educational improvement. This establishes a sustainable improvement cycle for improving the quality of education.

[0587] (Application Example 2)

[0588] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0589] In educational settings, it is difficult to appropriately assess learners' comprehension and concentration levels and provide real-time feedback. Traditional systems fail to adequately grasp learners' emotional states, making it difficult to improve teaching methods to suit individual learners. Furthermore, while there is a need to respond quickly to situations where learners are confused and provide appropriate educational materials, there is a lack of effective means to achieve this.

[0590] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0591] In this invention, the server includes means for acquiring information to analyze learners' awareness and concentration levels in an educational environment, means for evaluating learners' emotional states and the effectiveness of instructors' explanations using the acquired information, and means for notifying instructors in real time based on the evaluation and providing interactive educational materials to promote understanding among learners. This makes it possible to instantly grasp learners' emotional states and provide optimal teaching methods and materials for each individual learner.

[0592] The term "educational environment" refers to the place or situation in which learners acquire knowledge and skills, and includes physical classrooms and digital platforms.

[0593] The term "learner" refers to individuals who seek to acquire knowledge and skills within an educational environment.

[0594] "Recognition level" is an indicator that shows how well learners understand the educational content.

[0595] "Concentration" refers to the state in which learners are paying attention to educational activities.

[0596] "Emotional state" refers to the psychological state or reaction that learners exhibit during learning, and includes specific emotions such as "interesting" or "confusing."

[0597] "Interactive educational materials" refer to learning content and materials that allow learners to actively participate and interact with each other.

[0598] "Evaluation" refers to the process of analyzing learners' comprehension, concentration levels, and emotional states based on acquired data, and measuring the effectiveness of teaching methods.

[0599] "Notification" refers to the act of communicating information to instructors based on analysis results.

[0600] "Feedback" refers to information provided to improve teaching methods and enhance educational effectiveness, based on evaluation results, learner responses, and other factors.

[0601] The system implementing this invention analyzes learners' emotional states and comprehension levels in real time within an educational environment and provides appropriate feedback. The server uses cameras and microphones installed in the educational space to collect learners' facial expressions and voices in real time. This collected data is analyzed through emotion recognition software, such as the Microsoft Azure Emotion API, to detect the learners' emotional states.

[0602] The device receives analysis results from the server and notifies the instructor in real time with feedback tailored to the learner's emotional state. This notification includes specific advice for the instructor to adjust their teaching methods according to the learner's condition.

[0603] Furthermore, the server records changes in learners' comprehension and emotional states, and generates suggestions for evaluating long-term learning effectiveness. These suggestions include guidelines on how interactive educational materials should be used.

[0604] For example, in a physics lesson in a classroom, the server might detect that a student is experiencing "confusion" when faced with a difficult problem. The terminal would immediately notify the instructor with a message such as, "Some students are confused. Please try explaining using concrete examples." This notification allows the instructor to flexibly adjust their teaching methods.

[0605] Possible prompt statements for input to a generative AI model include the following:

[0606] "How can you engage students who aren't interested in historical discussions?"

[0607] "When we detect that students are confused while solving math problems, what kind of interactive learning materials can we provide?"

[0608] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0609] Step 1:

[0610] The server collects learners' facial expressions and audio data in real time through cameras and microphones installed in the educational space. Inputs include video streams and audio streams, which are stored in a database. The operations performed at this stage involve video capture by the cameras and audio recording by the microphones.

[0611] Step 2:

[0612] The server sends the collected facial expression and voice data to emotion recognition software (e.g., Microsoft Azure Emotion API) to analyze the learner's emotional state. The input is in the form of raw data, which the emotion recognition software processes to generate outputs representing the learner's intuitive psychological state, such as labels like "interesting" or "confused." Data preprocessing and emotion analysis are the main operations in this step.

[0613] Step 3:

[0614] The server inputs emotional state data obtained from emotion recognition software into an analysis device to evaluate learners' comprehension and concentration levels. At this stage, data calculations are performed to correlate emotional states with the progress of the lesson and generate evaluation results. The output provides individual learners' evaluations of comprehension and concentration levels.

[0615] Step 4:

[0616] The terminal notifies instructors of evaluation results in real time, helping them select appropriate teaching methods. The input is the evaluation results sent from the server, and the output is a specific feedback message presented to the instructor. The key feature here is that the notification system operates in real time.

[0617] Step 5:

[0618] The server stores evaluation results and emotional state records in a database and generates suggestions for improving teaching methods based on that data. Long-term learner data history is used as input, and actionable suggestions are generated as output. The main operations in this step are data analysis and suggestion generation.

[0619] Step 6:

[0620] The terminal provides the instructor with generated suggestions and directly offers learners interactive educational materials tailored to specific situations as needed. Input is data from the suggestion generator, and output is specific educational material. The operation here is an appropriate response based on the learner's situation.

[0621] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0622] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0623] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0624] [Fourth Embodiment]

[0625] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0626] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0627] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0628] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0629] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0630] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0631] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0632] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0633] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0634] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0635] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0636] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0637] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0638] This invention relates to a system that analyzes learners' comprehension and concentration levels in an educational environment and provides real-time feedback to instructors. This system consists of an information acquisition device, an analysis device, a notification device, a recording device, and a suggestion generation device.

[0639] Information acquisition process

[0640] The server collects learners' facial expressions, posture, and voice data via cameras and microphones installed in the classroom. This provides foundational data to understand the learners' current emotional state. For example, the camera detects facial movements when a learner frowns and records this as a change in facial expression.

[0641] Data analysis process

[0642] The server processes the collected data using advanced image and audio analysis algorithms. Image data is analyzed using facial recognition technology to determine whether the learner understands or is confused. Audio data is analyzed using audio analysis technology to extract characteristics of the instructor's speaking style (e.g., intonation and tempo) and evaluate the effectiveness of the instruction.

[0643] Real-time notification process

[0644] The device notifies the instructor in real time if it detects a decrease in the learner's concentration or that they are having difficulty understanding the material. For example, a message such as "Student D may be having difficulty understanding" might appear on the instructor's device. This information can then be used by the instructor to adjust their teaching methods on the spot.

[0645] Report generation process

[0646] After the lesson ends, the server automatically generates a comprehensive evaluation report based on the data collected. The report specifically shows which learners made progress in understanding and at what point their concentration wavered. This allows instructors to review the entire lesson and use the results to improve their teaching later.

[0647] Proposal generation process

[0648] The server is a device that uses recorded data and analysis results to generate improvement suggestions for instructors. This allows instructors to obtain specific, actionable directions for further deepening their teaching methods. For example, it might provide suggestions such as, "Incorporating visual aids into lessons may improve comprehension."

[0649] This system enables the integrated and dynamic assessment and response to learners' understanding and concentration within the educational environment, thereby promoting improvements in the quality of education.

[0650] The following describes the processing flow.

[0651] Step 1:

[0652] The server collects video and audio data in real time through cameras and microphones installed in the classroom. The cameras capture each student's face and expressions, and the microphones record the instructor's voice. During this process, the image data is processed frame by frame, and the audio data is converted to the appropriate format.

[0653] Step 2:

[0654] The server preprocesses the collected image and audio data. This preprocessing involves performing face recognition and expression extraction using the image data, and removing noise from the audio data to generate clear audio clips. This process prepares the data for analysis.

[0655] Step 3:

[0656] The server inputs preprocessed data into an AI model and performs analysis using a deep learning algorithm. The analysis classifies the learner's emotional state from their facial expressions and evaluates their concentration and comprehension levels. It also analyzes the instructor's speaking speed and intonation from audio data to determine if the instruction is effective.

[0657] Step 4:

[0658] The device receives analysis results from the server and sends real-time notifications to the instructor. These notifications include alerts if the learner's comprehension is declining or if their speaking style needs improvement. These notifications appear as pop-ups on the instructor's device.

[0659] Step 5:

[0660] At the end of each lesson, the server generates a report based on an overall assessment of comprehension and concentration levels derived from analysis. This report includes each learner's responses and changes in concentration at specific points during the lesson. Instructors can use this information to improve future lessons.

[0661] Step 6:

[0662] The server generates specific suggestions for improving teaching methods based on accumulated data. These suggestions include concrete actions to improve learner responses based on past data, supporting the growth of instructors.

[0663] (Example 1)

[0664] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0665] In educational settings, there is a lack of means to accurately grasp learners' levels of understanding and concentration in real time and to quickly convey that information to instructors. Furthermore, there is a need for effective methods to appropriately analyze changes in learners' emotional states and learning progress and use that information to improve instruction. This presents a challenge in that it is difficult for instructors to adjust their teaching methods to be optimal for each learner in a timely manner.

[0666] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0667] In this invention, the server includes means for acquiring the learner's emotional state, posture, and voice information in the learning environment; means for analyzing the learner's emotional state and the characteristics of the instructor's explanation using the image and voice information acquired by the means; and means for providing notifications that enable real-time adjustment of the teaching method based on the analysis results. This makes it possible to grasp the learner's level of understanding and concentration in real time and dynamically adjust the teaching method based on that.

[0668] The term "learning environment" refers to the place or space where education takes place, as well as the equipment and technology used within it.

[0669] A "learner" refers to an individual who participates in educational activities with the aim of acquiring knowledge and skills.

[0670] "Emotional state" refers to the psychological expressions and emotional states exhibited by learners, which can be observed through changes in facial expressions and voice.

[0671] "Postural information" refers to data related to the learner's body position and movement, and is used to determine their level of concentration and interest.

[0672] "Auditory information" refers to sound data collected within the learning environment, encompassing the speech of learners and instructors, as well as ambient sounds.

[0673] "Analysis results" refer to the conclusions and insights derived from processing the collected data.

[0674] "Notifications" refer to information and messages generated based on analysis and provided to instructors.

[0675] "Teaching methods" refer to the techniques and processes by which an instructor provides education to learners.

[0676] "Recording" refers to the act of saving data and information so that it can be referenced or analyzed later, or the result of doing so.

[0677] "Improvement suggestions" refer to specific proposals provided to improve the quality of education, based on recorded data and analysis.

[0678] This invention aims to implement a system that analyzes learners' comprehension and concentration levels in real time within educational settings and provides feedback to instructors. The details are described below.

[0679] The server utilizes cameras and microphones installed within the learning environment to collect learners' facial expressions, posture, and voice information. This data forms the basis for understanding the learners' emotional state at the time of acquisition.

[0680] Next, the server processes the collected data using advanced image and audio analysis software. Image data is analyzed based on facial expression recognition technology to identify how well the learner understands or is confused. Meanwhile, audio data is analyzed for changes in tone and tempo to evaluate the characteristics of the instructor's explanation.

[0681] The terminal notifies the instructor in real time based on the judgment results obtained from the server. This notification is displayed to the instructor in the form of, for example, "Student D may be having difficulty understanding." The instructor uses this information to adjust the lesson plan and approach on the spot.

[0682] Furthermore, after the lesson ends, the server automatically generates an evaluation report based on the analysis results. This report details which learners showed difficulties with understanding or concentration and at what point, and is used as feedback for instructors.

[0683] Furthermore, based on the recorded data and evaluation results, suggestions for improving teaching methods are generated. For example, a suggestion might be made such as, "Incorporating visual materials into lessons may improve learners' comprehension."

[0684] This system will promote increased efficiency and improved quality of instruction in educational settings.

[0685] Examples of prompt statements include the following:

[0686] "Could you provide an example of a notification message sent to an instructor if they determine that a student is losing focus during class?"

[0687] "Please provide three suggestions for improving learners' comprehension."

[0688] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0689] Step 1:

[0690] The server acquires the learner's facial expressions, posture, and voice information via the camera and microphone within the learning environment. It receives camera video and audio data as input and uses this to aggregate basic data about the learner's state in real time. Specifically, the image sensor captures facial expression data, and the voice sensor records the learner's speech and ambient sounds. The collected raw data is sent to the analysis module as output.

[0691] Step 2:

[0692] The server performs facial recognition using collected image data. It receives raw data as input and applies image processing algorithms to classify the learner's emotional state. For example, it determines states such as "understanding," "confused," and "indifferent" from eyebrow movements and changes in the mouth. As output, these analysis results are quantified as levels of concentration and comprehension and stored in a database.

[0693] Step 3:

[0694] The server performs speech analysis using audio data. It receives recorded audio as input and extracts speech features by applying a speech recognition algorithm. Specifically, it analyzes changes in the instructor's speaking tempo and volume and compares them with the learner's reactions. As output, an evaluation of the effectiveness of the instruction content is quantified and integrated with the results of facial expression analysis.

[0695] Step 4:

[0696] The device generates real-time notifications based on the analysis results. It receives analysis data as input and analyzes and makes decisions based on set thresholds. For example, if the "concentration level" falls below a predetermined value, it notifies the instructor that "Person A's concentration is declining." Actionable feedback is displayed on the device screen as output.

[0697] Step 5:

[0698] The server automatically generates evaluation reports based on the integrated data after each lesson. It receives the aggregated analytical data as input and processes it to visualize the progress of each learner's understanding and concentration levels. Specifically, the report outputs areas where understanding has improved and areas requiring improvement. The output is provided in a format viewable by instructors, which can be used to improve future lessons.

[0699] Step 6:

[0700] The server generates specific suggestions for improving teaching methods based on recorded data. Using analysis results stored in the database as input, it employs an AI model to derive optimal improvement suggestions. The output provides instructors with suggestions such as "increase the use of visual aids," which can be used to plan future lessons.

[0701] (Application Example 1)

[0702] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0703] In educational settings and at home, there is a growing need to dynamically grasp learners' comprehension and concentration levels and provide real-time feedback. However, traditional methods have faced challenges such as delayed feedback in educational settings and a lack of data necessary for providing appropriate guidance to individual learners. Furthermore, at home, it is currently difficult for parents to accurately grasp their children's learning progress and take appropriate action.

[0704] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0705] In this invention, the server includes means for acquiring information for analyzing the learner's level of understanding and concentration in an educational environment; analysis means for evaluating the learner's emotional state and the effectiveness of the educator's explanations to the learner using the information acquired by the means; and means for notifying the educator in real time based on the evaluation obtained by the analysis means. This enables the learner's condition to be grasped quickly and accurately, allowing educators and parents to take immediate action.

[0706] "Educational environment" is a general term for the places and circumstances in which educational activities are conducted for learners.

[0707] A "learner" refers to someone who receives education in order to acquire knowledge and skills.

[0708] "Comprehension level" is a measure that indicates the extent to which a learner understands the learning material.

[0709] "Concentration level" is a measure that indicates how much attention a learner is paying to a learning activity.

[0710] "Means of acquiring information" refers to devices and methods for collecting data about learners.

[0711] "Analysis means" refers to the processes and techniques used to analyze acquired information and evaluate the learner's state.

[0712] "Means of notification" refers to methods or devices for conveying specific information to educators or parents based on analysis results.

[0713] "Guardian" refers to an adult who is responsible for the upbringing of a learner.

[0714] "Means of generating suggestions" refers to methods and processes for creating suggestions that help improve teaching methods based on acquired and analyzed data.

[0715] This system aims to analyze learners' comprehension and concentration levels in educational and home environments and provide real-time feedback to educators and parents. The following describes its configuration and operating procedures in detail.

[0716] The server uses cameras and microphones installed in educational and home environments to acquire learners' facial expressions, posture, and voice data. This allows for the collection of data tailored to the situation in real time. The hardware uses a standard webcam as the camera and a directional microphone as the microphone.

[0717] The collected data is processed on the server. For image data, facial recognition is performed using OpenCV and TensorFlow. Audio data is analyzed using PyDub to evaluate the intonation and tempo of the speech. This analysis allows us to determine whether the learner understands or is confused.

[0718] Based on the analysis results, the device sends notifications to educators and parents. Using the Pushbullet API, messages are sent in real time to smartphones and tablets. For example, a notification might appear stating, "Learner A may be having difficulty understanding the material." This information can be used to immediately adjust teaching methods.

[0719] Furthermore, after a learning session ends, the server automatically generates a detailed report based on the collected and analyzed data. This report is provided to educators as a reference for improving their teaching methods and can also be used to support learning at home.

[0720] As a concrete example, when a student is working on a math assignment at home, facial recognition can detect signs of confusion, such as wrinkles between the eyebrows or changes in eye movement. The parent is then immediately notified that "Student A appears to be having difficulty understanding the material."

[0721] An example of a prompt to input into the generating AI model is: "Analyze my child's facial expressions while they are studying math, determine if they are having difficulty understanding, and send a notification."

[0722] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0723] Step 1:

[0724] The server collects learners' facial expressions, posture, and voice data in real time through installed cameras and microphones. This input data includes changes in facial expressions and vocal intonation, and is used as foundational data to understand the learners' emotional state. Specifically, the camera captures facial feature points, and the microphone records the tone and volume of the voice.

[0725] Step 2:

[0726] The server performs facial recognition processing on the acquired image data using OpenCV and TensorFlow. This process extracts feature points from the learner's face from the input data and analyzes what emotions the input facial image indicates. For example, if there are wrinkles between the eyebrows, it will be determined that the person is confused. The output will be the analysis result of the learner's emotional state.

[0727] Step 3:

[0728] The server uses PyDub to perform speech analysis on the audio data. This analysis examines the intonation and tempo of the input audio data to determine the learner's level of concentration and comprehension. Specifically, if the tone of voice suddenly becomes flat, it is determined that the learner's concentration may have been interrupted. The output of this analysis represents the results of the comprehension and concentration assessment.

[0729] Step 4:

[0730] Based on the analysis results, the device uses the Pushbullet API to send notifications to educators and parents. In this step, the device displays the analyzed results of the learner's comprehension and concentration levels as a message. For example, a notification saying "Learner A may be having difficulty understanding" is sent to the parent's smartphone. The input is the analysis results, and the output is the notification message.

[0731] Step 5:

[0732] The server automatically generates a comprehensive evaluation report based on the data collected after each learning session. This report includes information on when the learner understood the material or when their concentration wavered, providing valuable insights for improving future teaching methods. Specifically, the server organizes the collected data chronologically and summarizes the evaluation results. The input is the previously analyzed data, and the output is the evaluation report.

[0733] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0734] This invention combines an emotion engine with a system for analyzing learners' comprehension and concentration levels in an educational environment. This system consists of an information acquisition device, an analysis device, a notification device, a recording device, a suggestion generation device, and an emotion engine.

[0735] The process of information acquisition and emotion recognition

[0736] The server collects learners' facial expressions and voice data in real time through cameras and microphones installed in the classroom. The emotion engine uses this data to recognize the learners' emotional state and identify emotions such as "interesting" or "confusing."

[0737] Data analysis and coaching feedback process

[0738] The server inputs data, including emotional states recognized by the emotion engine, into the analysis device. The analysis device evaluates the relationship between the learner's emotional state, comprehension level, and concentration level, and makes a comprehensive judgment on the effectiveness of the instruction. As a result, it is possible to understand how the learner is receiving the lesson content. For example, if a learner is "confused," it is determined that the teaching method at that point should be reconsidered.

[0739] Real-time notifications and the process of improving instruction

[0740] The device provides real-time feedback to instructors based on the analysis results. For example, a message such as "Ms. E seems a little confused" might appear on the instructor's device, allowing the instructor to adjust their teaching method immediately. Feedback is also provided if certain emotions are repeatedly observed.

[0741] Recording and proposal generation process

[0742] The server accumulates past emotional data and analysis results to evaluate changes in learners' states over the long term. Based on this, it generates suggestions for optimizing teaching methods. For example, if it is confirmed that some learners' comprehension improves with audiovisual stimuli, the use of visual materials will be suggested.

[0743] This system enables the improvement of educational quality by conducting multifaceted analysis, including learners' emotions. Instructors can grasp learners' emotions in real time and gain concrete means to implement appropriate instruction.

[0744] The following describes the processing flow.

[0745] Step 1:

[0746] The server uses cameras and microphones in the classroom to collect learners' facial expressions and audio data in real time. The cameras capture each learner's facial expressions, and the microphones record the tone and intonation of their individual voices. This data is temporarily stored in a database in its raw state.

[0747] Step 2:

[0748] The server preprocesses the collected data. Image data undergoes image filtering to identify facial features and extract expressions. Audio data is processed through noise filtering and audio frame extraction, and then formatted to be suitable for audio analysis methods.

[0749] Step 3:

[0750] The server supplies pre-processed data to the emotion engine. The emotion engine uses machine learning algorithms to analyze emotions from facial expressions and voice, determining emotional states such as "joy," "confusion," and "distraction." This enables instantaneous diagnosis of the learner's emotions.

[0751] Step 4:

[0752] The device receives sentiment analysis results sent from the server and notifies the instructor. Specifically, it displays a message such as, "Student F is confused by the lesson content," prompting the instructor to adjust the lesson's pace and explanation methods. This notification enables flexible responses during the lesson.

[0753] Step 5:

[0754] The server tracks and stores emotional data and associated analytical information over long periods. The stored data is organized in a format that allows for the evaluation of emotional changes over time. As a result, instructors can understand the temporal evolution of learners' emotions and use this information to improve teaching methods.

[0755] Step 6:

[0756] The server generates specific suggestions for optimizing teaching methods based on the combined data. These suggestions include actionable improvements, such as "Increasing visual aids can improve comprehension," providing practical guidelines for instructors.

[0757] (Example 2)

[0758] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0759] In educational settings, it is difficult to instantly grasp learners' levels of understanding and concentration and provide appropriate instruction; therefore, efficient methods are needed to maximize learning effectiveness. In particular, there is a need to develop a system that allows instructors to respond quickly by providing real-time feedback on learners' emotional states.

[0760] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0761] In this invention, the server includes means for collecting video and audio from the environment to identify the learner's emotions; means for analyzing the collected video and audio and evaluating the learner's emotions and learning effectiveness using a specified algorithm; and means for providing real-time feedback to the instructor according to the evaluated emotional state. This makes it possible to instantly grasp the learner's level of understanding and emotional state and improve the quality of instruction.

[0762] A "learner" is someone who seeks to acquire knowledge and skills in an educational environment.

[0763] "Emotions" refer to the mental state of a learner and include states such as interest, confusion, happiness, and anxiety.

[0764] "Video" refers to dynamic or static image data that visually represents the learner's facial expressions and actions.

[0765] "Speech" refers to acoustic signals, including the learner's utterances and tone of voice.

[0766] An "algorithm" is a set of procedures or computational steps for solving a specific problem, and it is used as a means of analyzing learners' emotions.

[0767] "Feedback" is information provided to instructors regarding the learner's current state, which allows them to adjust their teaching methods accordingly.

[0768] A "threshold" is a specific numerical value or level set as a standard for evaluating a learner's level of understanding and concentration.

[0769] A "suggestion" is information that provides instructors with specific guidance on improving and optimizing teaching methods, based on recorded data and analysis results.

[0770] This invention is a system for analyzing learners' comprehension and concentration levels in an educational environment and optimizing teaching methods. The system consists of three main components: a server, terminals, and users.

[0771] First, the server uses cameras and microphones placed in the classroom to collect learners' facial expressions and voice data in real time. This hardware operates continuously to facilitate data acquisition. The collected data is processed by an emotion engine to analyze the learners' emotional state. This emotion engine is equipped with emotion recognition algorithms and has advanced capabilities to analyze subtle changes in facial expressions and tone of voice.

[0772] Next, based on this analyzed data, the device provides real-time feedback to the instructor. The device immediately notifies the instructor according to the analysis results, displaying specific messages such as, "Student E is confused." This allows the instructor to immediately adopt teaching methods that respond to the learner's situation.

[0773] Furthermore, the server accumulates this data and analysis results over the long term, continuously evaluating changes in learners' situations. Based on this long-term data, the server can generate specific suggestions for improving teaching methods. This allows instructors to obtain valuable information for continuously optimizing their teaching methods.

[0774] For example, if data reveals that visual learning materials improve the comprehension of certain learners, then teaching methods can be proposed. Instructors can then use these suggestions to further improve the quality of their instruction.

[0775] An example of a prompt for a generative AI model is: "Please explain how to analyze learners' emotions in a classroom and evaluate their comprehension and concentration levels. Also, show how this data can be used to optimize instruction."

[0776] This system enables data-driven instruction to improve the quality of education.

[0777] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0778] Step 1:

[0779] The server uses cameras and microphones installed in the classroom to collect learners' facial expressions and voice data as input. Specifically, the cameras capture video at several frames per second and send it to the server as digital data. The microphones record learners' speech and ambient sounds in real time and send them to the server in digital format. This generates a petabyte-scale dataset that reflects the learners' state.

[0780] Step 2:

[0781] The server takes the collected facial expression and audio data as input and entrusts the processing to the emotion engine. Specifically, the emotion engine uses a facial recognition algorithm to analyze the facial features of each frame and uses speech recognition technology to evaluate the tone of voice and emotion. As output, the server generates emotion states labeled as "interesting" or "confused." These labels indicate the learner's state in real time.

[0782] Step 3:

[0783] The server sends emotional state data obtained from the emotion engine as input to the analysis device. The analysis device uses data mining techniques to compare this data with past data and evaluate the learner's level of understanding and concentration. As output, evaluation results regarding the effectiveness of specific teaching methods are generated. This prepares the server to receive feedback on the effectiveness of the instruction as numerical values ​​and evaluation comments.

[0784] Step 4:

[0785] The terminal receives evaluation results from the server as input and provides real-time feedback to the instructor. Specifically, the terminal displays a warning message such as "Person E is confused," allowing the instructor to immediately adjust their teaching methods based on the feedback. The output is information provided to the instructor and improvements to the teaching based on that information. This rapid feedback enables instructors to respond immediately to provide appropriate instruction.

[0786] Step 5:

[0787] The server records all data and analysis results over the long term as input. Specifically, it stores data in a database and organizes it to be useful for subsequent analysis. Based on this data, a suggestion generator generates suggestions for optimizing teaching methods. As output, instructors are provided with concrete suggestions for long-term educational improvement. This establishes a sustainable improvement cycle for improving the quality of education.

[0788] (Application Example 2)

[0789] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0790] In educational settings, it is difficult to appropriately assess learners' comprehension and concentration levels and provide real-time feedback. Traditional systems fail to adequately grasp learners' emotional states, making it difficult to improve teaching methods to suit individual learners. Furthermore, while there is a need to respond quickly to situations where learners are confused and provide appropriate educational materials, there is a lack of effective means to achieve this.

[0791] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0792] In this invention, the server includes means for acquiring information to analyze learners' awareness and concentration levels in an educational environment, means for evaluating learners' emotional states and the effectiveness of instructors' explanations using the acquired information, and means for notifying instructors in real time based on the evaluation and providing interactive educational materials to promote understanding among learners. This makes it possible to instantly grasp learners' emotional states and provide optimal teaching methods and materials for each individual learner.

[0793] The term "educational environment" refers to the place or situation in which learners acquire knowledge and skills, and includes physical classrooms and digital platforms.

[0794] The term "learner" refers to individuals who seek to acquire knowledge and skills within an educational environment.

[0795] "Recognition level" is an indicator that shows how well learners understand the educational content.

[0796] "Concentration" refers to the state in which learners are paying attention to educational activities.

[0797] "Emotional state" refers to the psychological state or reaction that learners exhibit during learning, and includes specific emotions such as "interesting" or "confusing."

[0798] "Interactive educational materials" refer to learning content and materials that allow learners to actively participate and interact with each other.

[0799] "Evaluation" refers to the process of analyzing learners' comprehension, concentration levels, and emotional states based on acquired data, and measuring the effectiveness of teaching methods.

[0800] "Notification" refers to the act of communicating information to instructors based on analysis results.

[0801] "Feedback" refers to information provided to improve teaching methods and enhance educational effectiveness, based on evaluation results, learner responses, and other factors.

[0802] The system implementing this invention analyzes learners' emotional states and comprehension levels in real time within an educational environment and provides appropriate feedback. The server uses cameras and microphones installed in the educational space to collect learners' facial expressions and voices in real time. This collected data is analyzed through emotion recognition software, such as the Microsoft Azure Emotion API, to detect the learners' emotional states.

[0803] The device receives analysis results from the server and notifies the instructor in real time with feedback tailored to the learner's emotional state. This notification includes specific advice for the instructor to adjust their teaching methods according to the learner's condition.

[0804] Furthermore, the server records changes in learners' comprehension and emotional states, and generates suggestions for evaluating long-term learning effectiveness. These suggestions include guidelines on how interactive educational materials should be used.

[0805] For example, in a physics lesson in a classroom, the server might detect that a student is experiencing "confusion" when faced with a difficult problem. The terminal would immediately notify the instructor with a message such as, "Some students are confused. Please try explaining using concrete examples." This notification allows the instructor to flexibly adjust their teaching methods.

[0806] Possible prompt statements for input to a generative AI model include the following:

[0807] "How can you engage students who aren't interested in historical discussions?"

[0808] "When we detect that students are confused while solving math problems, what kind of interactive learning materials can we provide?"

[0809] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0810] Step 1:

[0811] The server collects learners' facial expressions and audio data in real time through cameras and microphones installed in the educational space. Inputs include video streams and audio streams, which are stored in a database. The operations performed at this stage involve video capture by the cameras and audio recording by the microphones.

[0812] Step 2:

[0813] The server sends the collected facial expression and voice data to emotion recognition software (e.g., Microsoft Azure Emotion API) to analyze the learner's emotional state. The input is in the form of raw data, which the emotion recognition software processes to generate outputs representing the learner's intuitive psychological state, such as labels like "interesting" or "confused." Data preprocessing and emotion analysis are the main operations in this step.

[0814] Step 3:

[0815] The server inputs emotional state data obtained from emotion recognition software into an analysis device to evaluate learners' comprehension and concentration levels. At this stage, data calculations are performed to correlate emotional states with the progress of the lesson and generate evaluation results. The output provides individual learners' evaluations of comprehension and concentration levels.

[0816] Step 4:

[0817] The terminal notifies instructors of evaluation results in real time, helping them select appropriate teaching methods. The input is the evaluation results sent from the server, and the output is a specific feedback message presented to the instructor. The key feature here is that the notification system operates in real time.

[0818] Step 5:

[0819] The server stores evaluation results and emotional state records in a database and generates suggestions for improving teaching methods based on that data. Long-term learner data history is used as input, and actionable suggestions are generated as output. The main operations in this step are data analysis and suggestion generation.

[0820] Step 6:

[0821] The terminal provides the instructor with generated suggestions and directly offers learners interactive educational materials tailored to specific situations as needed. Input is data from the suggestion generator, and output is specific educational material. The operation here is an appropriate response based on the learner's situation.

[0822] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0823] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0824] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0825] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0826] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0827] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0828] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0829] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0830] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0831] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0832] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0833] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0834] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0835] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0836] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0837] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0838] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0839] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0840] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0841] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0842] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0843] The following is further disclosed regarding the embodiments described above.

[0844] (Claim 1)

[0845] A device for acquiring information to analyze learners' comprehension and concentration levels in an educational environment,

[0846] An analysis device that uses the information acquired by the aforementioned device to evaluate the learner's emotional state and the effectiveness of the instructor's explanation to the learner,

[0847] A device that notifies the instructor in real time based on the evaluation obtained by the aforementioned analysis device,

[0848] A device that records the aforementioned evaluation and notification content and generates suggestions for improving teaching methods,

[0849] A system that includes this.

[0850] (Claim 2)

[0851] The system according to claim 1, wherein the device for notifying the instructor is controlled to notify the instructor when it detects that the learner's level of understanding has fallen below a pre-set threshold.

[0852] (Claim 3)

[0853] The system according to claim 1, wherein the device that generates the proposal makes the proposal for improving the teaching method based on the temporal progression of the learner's concentration level.

[0854] "Example 1"

[0855] (Claim 1)

[0856] In a learning environment, means of acquiring learners' emotional state, posture, and auditory information,

[0857] A means for analyzing the learner's emotional state and the characteristics of the instructor's explanation using the image and audio information acquired by the aforementioned means,

[0858] A means of providing notifications that enable real-time adjustment of teaching methods based on analysis results,

[0859] A means for recording the aforementioned analysis results and notification information, and for creating and providing an overall evaluation report of the course based on the acquired data,

[0860] A means of generating suggestions for improving teaching methods based on evaluated and recorded information,

[0861] A system that includes this.

[0862] (Claim 2)

[0863] The system according to claim 1, which provides real-time notification when a learner's level of understanding falls below a set threshold.

[0864] (Claim 3)

[0865] The system according to claim 1, which provides the aforementioned improvement suggestions based on the temporal changes in the learner's level of understanding and concentration.

[0866] "Application Example 1"

[0867] (Claim 1)

[0868] In an educational environment, means of acquiring information to analyze learners' level of understanding and concentration,

[0869] An analytical means for evaluating the learner's emotional state and the effectiveness of the educator's explanation to the learner, using the information obtained by the aforementioned means,

[0870] Based on the evaluation obtained by the aforementioned analysis means, a means for notifying educators in real time,

[0871] A means for recording the aforementioned evaluation and notification content and generating suggestions for improving teaching methods,

[0872] A means of analyzing the learner's level of understanding and concentration at home and notifying parents,

[0873] A system that includes this.

[0874] (Claim 2)

[0875] The system according to claim 1, wherein the means for notifying the educator or guardian is controlled to notify the educator or guardian when it detects that the learner's level of understanding has fallen below a pre-set threshold.

[0876] (Claim 3)

[0877] The system according to claim 1, wherein the means for generating the proposal makes a proposal for improving the teaching method based on the temporal progression of the learner's concentration level.

[0878] "Example 2 of combining an emotion engine"

[0879] (Claim 1)

[0880] A device for collecting images and sounds from the environment in order to identify the learner's emotions,

[0881] A means for analyzing the collected video and audio and evaluating the learner's emotions and learning effects using a specified algorithm,

[0882] A means of providing real-time feedback to the instructor in accordance with the aforementioned evaluated emotional state,

[0883] A means for recording the aforementioned emotional state and evaluation results over the long term and generating suggestions for improving teaching methods,

[0884] A system that includes this.

[0885] (Claim 2)

[0886] The system according to claim 1, wherein the means for providing feedback to the instructor is controlled to notify the instructor when it detects that the learner's level of understanding has fallen below a set threshold.

[0887] (Claim 3)

[0888] The system according to claim 1, wherein the means for generating the aforementioned proposal makes suggestions for improving teaching methods based on the temporal progression of the learner's level of concentration.

[0889] "Application example 2 when combining with an emotional engine"

[0890] (Claim 1)

[0891] A device for acquiring information to analyze learners' level of awareness and concentration in an educational environment,

[0892] An analysis device that uses the information acquired by the aforementioned device to evaluate the learner's emotional state and the effectiveness of the instructor's explanation to the learner,

[0893] A device that notifies the instructor in real time based on the evaluation obtained by the aforementioned analysis device,

[0894] A device that records the aforementioned evaluation and notification content and generates suggestions for improving teaching methods,

[0895] A device that provides interactive educational materials to encourage understanding in learners when their emotional state falls below a target value,

[0896] A system that includes this.

[0897] (Claim 2)

[0898] The system according to claim 1, wherein the device for notifying the instructor is controlled to notify the instructor when it detects that the learner's level of recognition has fallen below a pre-set standard, and is also controlled to automatically provide interactive educational materials to the learner.

[0899] (Claim 3)

[0900] The system according to claim 1, wherein the device that generates the proposal makes suggestions for improving the teaching method based on changes in the learner's concentration level, and generates feedback to verify and optimize the effectiveness of the interactive educational materials provided to the learner. [Explanation of Symbols]

[0901] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. In an educational environment, means of acquiring information to analyze learners' level of understanding and concentration, An analytical means for evaluating the learner's emotional state and the effectiveness of the educator's explanation to the learner, using the information obtained by the aforementioned means, Based on the evaluation obtained by the aforementioned analysis means, a means for notifying educators in real time, A means for recording the aforementioned evaluation and notification content and generating suggestions for improving teaching methods, A means of analyzing the learner's level of understanding and concentration at home and notifying parents, A system that includes this.

2. The system according to claim 1, wherein the means for notifying the educator or guardian is controlled to notify the educator or guardian when it detects that the learner's level of understanding has fallen below a pre-set threshold.

3. The system according to claim 1, wherein the means for generating the proposal makes a proposal for improving the teaching method based on the temporal progression of the learner's concentration level.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A