system
The system addresses teacher challenges by using voice and video acquisition, processing, and content generation to provide immediate, personalized learning support, improving educational quality and efficiency.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-21
- Publication Date
- 2026-05-07
AI Technical Summary
Teachers face challenges in accurately assessing students' learning progress and understanding levels, responding to individual questions, detecting mental health issues, and managing abnormal behaviors, leading to increased workload and reduced educational quality and efficiency.
A system that incorporates voice and video acquisition, speech and natural language processing, video analysis, content generation, and anomaly detection to provide immediate responses and personalized learning content based on students' needs, reducing teacher burden and improving instruction quality.
The system enables efficient, individualized instruction by automating responses to student questions, detecting abnormal behaviors, and optimizing learning content, thereby reducing teacher workload and enhancing educational quality and efficiency.
Smart Images

Figure 2026074961000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the conventional educational field, there has been a problem that it is a great burden for teachers to appropriately grasp the learning progress and understanding level of each student, and in addition, it is difficult to quickly respond to the questions of individual students. Also, it has been difficult to early detect the mental health and abnormal behaviors of students and take appropriate measures. These problems have also caused long working hours and mental stress for teachers, and have reduced the quality and efficiency of education.
Means for Solving the Problems
[0005] This invention provides a system that enables immediate responses to questions by incorporating a voice acquisition means that acquires student voice information and converts it into text data, and a natural language processing means that analyzes the text data to identify the content of the student's question. Furthermore, it includes a video acquisition means and a video analysis means that acquire student video information and evaluate their learning status and emotional state, providing a function to detect abnormal behavior in advance, thereby realizing a system that reduces the burden on teachers while providing learning content optimized for each individual student. In addition, by providing the generated learning content to students and accumulating the feedback data, the quality of individual instruction is improved and the system is optimized.
[0006] "Voice acquisition means" refers to devices or methods that capture voice information from students and record it as data.
[0007] "Speech recognition means" refers to a technology or system that analyzes acquired speech information and converts it into text data.
[0008] "Natural language processing techniques" refer to technologies that analyze the structure and meaning of language in order to identify the content of students' questions from text data.
[0009] "Means of acquiring video footage" refers to devices and technologies that capture and collect video information, including students' facial expressions and movements.
[0010] "Video analysis means" refers to technology that analyzes video information of students to evaluate their learning progress and emotional state.
[0011] "Content generation means" refers to technologies and methods for creating individualized learning content based on students' learning history and analysis results.
[0012] "Content delivery means" refers to the systems and methods used to provide generated learning content to students.
[0013] "Learning methods" refer to technologies that accumulate feedback data and new learning data to effectively optimize the entire system.
[0014] "Anomaly detection means" refers to technologies and methods for detecting student malaise or abnormal behavior in advance and generating alerts.
[0015] "Real-time response methods" refer to technologies and methods that provide immediate answers to students' questions. [Brief explanation of the drawing]
[0016] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12]It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined.
Mode for Carrying Out the Invention
[0017] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0018] First, the language used in the following description will be explained.
[0019] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0020] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0021] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0022] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0024] [First Embodiment]
[0025] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0026] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0029] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0032] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0036] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0037] This invention is a learning support system using an AI robot teacher in educational settings, aiming to reduce the burden on teachers and improve the quality of individualized instruction for students. This system mainly consists of terminals and a server.
[0038] The terminals are installed in the classrooms and are equipped with microphones to capture student voices in real time, and cameras to monitor students' facial expressions and behavior. As a result, audio and video information is acquired simultaneously and transmitted to the server.
[0039] The server uses a speech recognition engine to convert the audio information sent from the terminal into text data. At this stage, questions and requests are extracted from the audio. Furthermore, natural language processing is used to analyze the text data and understand the intent behind the student's questions. This analysis generates appropriate answers and learning content.
[0040] The server uses deep learning algorithms to analyze the video data, analyzing students' emotional states and concentration levels. This allows teachers to identify students who require individual attention and to evaluate the overall learning environment of the class. Furthermore, based on this data, an anomaly detection system immediately generates an alert if it detects any student distress or abnormal behavior, notifying the teacher.
[0041] Furthermore, the server uses a generative AI model to generate personalized learning content based on students' past learning history and progress data. This content is adjusted in difficulty and content according to the student's level of understanding, providing optimal learning materials tailored to individual needs.
[0042] Teachers, as users, can utilize the data and reports provided by their devices to understand their students' situations and develop effective teaching plans. This frees teachers from the detailed responses required in daily lessons, allowing them to dedicate more time and effort to individualized instruction.
[0043] Thus, the AI robot teacher system of the present invention can contribute to improving the efficiency and quality of educational settings by utilizing speech recognition and video analysis to automate student learning support.
[0044] The following describes the processing flow.
[0045] Step 1:
[0046] The device uses a microphone to capture audio information emitted by students in the classroom, while simultaneously capturing students' facial expressions and movements with a camera. This allows for the collection of audio and video data in real time.
[0047] Step 2:
[0048] The device sends the collected audio data to the server. The server receives this data and converts it into text data using a speech recognition engine. The recognized text is then processed to facilitate smooth responses to the teacher's questions.
[0049] Step 3:
[0050] The server analyzes the converted text data using natural language processing techniques to understand the intent and content of the questions. Based on this analysis, it generates answers to help students understand the questions.
[0051] Step 4:
[0052] The terminal sends video data to the server for analysis. The server uses deep learning technology to analyze the students' facial expressions and movements to evaluate their emotional state and level of concentration.
[0053] Step 5:
[0054] The server uses anomaly detection methods based on the analysis results to detect student malaise or abnormal behavior. If detected, an alert is immediately generated and notified to the user.
[0055] Step 6:
[0056] The server references the student's past learning history and generates optimized, personalized learning content using a generative AI model. The generated content is customized to the student's current level of understanding and progress.
[0057] Step 7:
[0058] The device provides students with generated learning content and receives real-time feedback from students through the device. This feedback is used to generate the next learning content.
[0059] Step 8:
[0060] Users can use real-time data and reports provided by their devices to understand students' learning progress and plan effective instruction. Teachers can more easily provide instruction based on quantitative data.
[0061] (Example 1)
[0062] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0063] There is a need to reduce the burden on teachers in educational settings and improve the quality of individualized instruction for students. Traditional methods have made it difficult to effectively respond to students' learning situations and individual needs, and it has also been difficult to grasp the condition of individual students in large groups. Furthermore, there are insufficient means to detect and respond to student distress or abnormal behavior at an early stage, so these challenges need to be addressed.
[0064] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0065] In this invention, the server includes a recording device for acquiring audio, a conversion device for converting audio information into text, and an analysis device for analyzing text information and extracting intent. This enables effective individualized instruction tailored to the student's learning situation and needs.
[0066] A "recording device" is a device used to collect audio information from students.
[0067] A "conversion device" is a device used to convert acquired audio information into text data.
[0068] An "analysis device" is a device that examines text data in detail to extract the student's intentions and the content of their questions.
[0069] A "filming device" is a device used to acquire video information of students.
[0070] A "judgment device" is a device that analyzes video information of students to evaluate their learning progress and emotional state.
[0071] A "generation device" is a device that creates individualized learning materials based on students' learning history and analysis results.
[0072] A "distribution device" is a device used to supply students with the generated teaching materials.
[0073] A "memory device" is a device used to store learning data and optimize the system.
[0074] A "detection device" is a device that detects students' poor health or abnormal behavior and generates alerts in advance.
[0075] A "response device" is a device that provides immediate answers to students' questions.
[0076] This invention is an AI system intended to support classroom instruction in educational settings. It mainly consists of a terminal installed in the classroom and a server for data processing. The terminal includes a high-sensitivity microphone for acquiring student voice information and a high-resolution camera for recording student video information.
[0077] The server uses a speech recognition engine to convert the audio information sent from the terminal into text data. A common example of software used in this speech recognition process is a cloud-based speech recognition service. The text data is then analyzed using natural language processing techniques to clarify the intent behind the student's questions and requests. For example, a natural language processing library might be used for this process.
[0078] Regarding video data, the server utilizes deep learning to evaluate students' emotional states and concentration levels. The algorithms used in this evaluation process include open-source machine learning frameworks. Furthermore, if any distress or abnormal behavior is detected based on the analysis results, an alert is immediately generated and notified to the teacher's terminal.
[0079] Furthermore, the server analyzes students' past learning history and uses a generative AI model to generate personalized learning materials. These materials are adjusted according to the student's level of understanding, providing a learning experience tailored to individual learning needs.
[0080] For example, if a student says, "I don't understand the meaning of differentiation," the system can analyze their intent, generate and provide learning materials that include the basic concepts of differentiation. An example of a prompt from this system is: "When a student asks a question about something they don't understand during class, analyze their intent and generate appropriate learning content."
[0081] As described above, the present invention reduces the burden on teachers and makes it possible to efficiently provide individualized instruction tailored to each student.
[0082] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0083] Step 1:
[0084] The device uses a microphone to capture audio from students in the classroom. The input is the student's voice, which is received by the device. Specifically, during class, asking a question like, "How do you solve this equation?" collects the audio data. The output is the collected audio data, which is sent to the server.
[0085] Step 2:
[0086] The server processes the transmitted audio data using a speech recognition engine and converts it into text data. The audio data received from the terminal is used as input. Specifically, the speech recognition engine recognizes the audio "How do you solve this equation?" and generates the corresponding text "How do you solve this equation?". The generated text data is obtained as output.
[0087] Step 3:
[0088] The server analyzes the generated text data using natural language processing techniques to extract the intent behind the student's question. Text data obtained from speech recognition is used as input. Specifically, the natural language processing technique identifies the intent, "I want to know how to solve this equation." The output reveals the intent behind the question.
[0089] Step 4:
[0090] The server processes video data transmitted from the terminal using a deep learning algorithm for emotion analysis and concentration assessment. The input is the video data received from the terminal. Specifically, it analyzes the video and detects an emotional state such as "confused" from the student's facial expressions and posture. The output is the analyzed emotional state information.
[0091] Step 5:
[0092] The server uses a generative AI model to generate learning content tailored to each student, based on the intent and emotional state of the question. The input includes the student's past learning history and analyzed data. Specifically, the AI model automatically creates materials including "basic equations and how to solve them." The output is personalized learning content.
[0093] Step 6:
[0094] The user, the teacher, reviews the generated learning content and analysis results, and provides feedback to students based on them. The inputs include content provided by the server and student status information. Specifically, the teacher utilizes the information from the system to provide individualized instruction to students during class. The output is the provision of appropriate guidance and support to students.
[0095] (Application Example 1)
[0096] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0097] To improve factory production efficiency, it is necessary to accurately understand workers' actions and emotional states and provide them with the most appropriate work procedures in a timely manner. However, conventional systems have difficulty analyzing workers' actions and emotions in real time, and are unable to detect potential hazards in advance and respond appropriately, resulting in insufficient improvements in work efficiency and safety.
[0098] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0099] In this invention, the server includes voice acquisition means for acquiring worker movement information, video analysis means for analyzing worker video information and evaluating production efficiency, and anomaly detection means for alerting inappropriate movements in advance. This makes it possible to grasp the worker's movements and emotions in real time, provide optimal work procedures, and avoid potential hazards.
[0100] "Workers" refer to people who assemble parts or process products within a factory, and their actions and emotional states are the targets for management and optimization.
[0101] "Motion information" refers to data about the movements and work attitude of workers, and is real-time information acquired through sensors such as cameras.
[0102] A "voice acquisition means" is a device that has the function of collecting the voice emitted by a worker as digital data using a microphone or other device.
[0103] "Speech recognition means" refers to a technology that has the function of converting acquired speech data into text data and analyzes the content of what the worker says as text information.
[0104] "Natural language processing means" refers to technologies that analyze text data generated by speech recognition means and understand its context and intent.
[0105] "Video acquisition means" refers to a device that has the function of recording the actions of workers as video using a camera or similar device, and is used for analyzing their movements and emotions.
[0106] "Video analysis means" refers to technology that processes acquired video data to evaluate the productivity and emotional state of workers.
[0107] A "content generation method" is a system that has the function of generating individually optimized work content based on the results of analysis of the worker's actions and emotions.
[0108] A "content delivery means" is a device that has the function of communicating the generated optimization work content to workers and supporting them in carrying out their tasks.
[0109] "Learning methods" refer to technologies that accumulate feedback data and newly acquired work data to improve the accuracy and efficiency of the system.
[0110] An "anomaly detection method" is a technology that detects inappropriate actions or potential hazards by workers in real time and issues an immediate alert.
[0111] A "real-time response system" is a technology that provides optimal work procedures and advice in response to the worker's instructions and situation.
[0112] To implement this invention, a system is configured in which a server, terminal, and user cooperate to evaluate and optimize the worker's actions in real time. The server continuously acquires data from the work site by utilizing voice acquisition means and video acquisition means to collect worker action information. The voice acquisition means uses a microphone to capture the worker's speech and converts it into text data using speech recognition means. Natural language processing software analyzes the worker's intentions and instructions from the text data.
[0113] Simultaneously, the video data collected by the video acquisition system is processed using video analysis tools that apply deep learning algorithms to evaluate the workers' actions and emotions. This makes it possible to grasp production efficiency and the workers' psychological state in real time.
[0114] Based on the analysis results, the server uses a generation AI model to create individually optimized work instructions in the content generation means, and provides appropriate work instructions to the user (worker) via the content provision means. In addition, if the anomaly detection means detects inappropriate operation or potential danger, an alert is immediately generated and notified to the worker and supervisor.
[0115] The learning method analyzes accumulated feedback data and newly acquired work data to continuously optimize the system. The hardware used in this process includes professional cameras, high-sensitivity microphones, and a high-performance computing platform located on a server as computing resources. Software such as OpenCV, Keras, and SpeechRecognition are employed.
[0116] As a concrete example, when a worker is assembling parts, the server detects the worker's anxious expression using video analysis, and a generative AI model generates encouraging comments and provides voice feedback through a content delivery system, such as "Please feel free to contact us if you have any problems." An example of a prompt message would be, "Please propose a new strategy to optimize the next automation process." In this way, the system plays a role in enhancing the smoothness and safety of work.
[0117] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0118] Step 1:
[0119] The server receives voice data from the worker, acquired via microphone, as input from the terminal. This voice data is then used as input for speech recognition, digital signal processing is performed, and it is converted into text data for output. Specifically, the characteristics of the voice waveform are analyzed, and recognition is performed using a language model.
[0120] Step 2:
[0121] The server uses the text data output in step 1 as input and processes it with natural language processing. Here, text analysis is performed to identify the intent behind the worker's instructions and questions, and data is output to determine actions based on that intent. Context and intent are interpreted using natural language processing technology.
[0122] Step 3:
[0123] The server acquires video information of workers as input via cameras connected to terminals. This data is then processed by a video analysis system, and a deep learning algorithm is used to evaluate production efficiency and the emotional state of the workers, outputting the evaluation results. Facial expressions and movement patterns are analyzed from the video data to infer psychological state and work speed.
[0124] Step 4:
[0125] The server uses the intent and evaluation results obtained in steps 2 and 3 as input to generate individually optimized work content using a generative AI model. This generated work content is output by a content generation means. At this stage, prompts are used to generate work procedures and advice as specific instructions.
[0126] Step 5:
[0127] The server provides the content generated in step 4 to the user (worker) via the content delivery means. The outputted work instructions and advice are transmitted to the worker as audio or video, supporting the worker in properly performing their tasks.
[0128] Step 6:
[0129] The server stores feedback data collected during the work process and newly acquired work data as input to its learning mechanism. This allows for long-term data analysis to optimize the system and improve its accuracy. Statistical analysis of the data and machine learning are used to improve the accuracy of future work instructions.
[0130] Step 7:
[0131] The server uses anomaly detection mechanisms to detect inappropriate worker behavior and potential hazards in real time. Based on this input, it immediately generates and outputs alerts, notifying the user (worker) and relevant supervisors. This ensures workplace safety and enables a rapid response.
[0132] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0133] This invention aims to achieve more effective learning support by combining an emotion engine with a system that utilizes AI robot teachers in educational settings. This system mainly consists of a terminal, a server, and emotion recognition means.
[0134] The terminal has the ability to acquire students' audio and video in real time within the classroom and transmits this data to a server. The server converts the received audio data into text data using speech recognition technology and analyzes the content of the questions using natural language processing technology. Based on these results, specific answers and individualized learning content are generated to support students' learning.
[0135] The server analyzes the video data using deep learning technology to evaluate students' emotional states and concentration levels based on their facial expressions and movements. This evaluation allows the system to track students' learning progress, and individualized support is provided as needed through content generation methods.
[0136] The emotion engine is used specifically to analyze the emotional state of the user, the teacher. This allows for feedback that helps maintain or improve the quality of instruction based on the teacher's own emotions.
[0137] Specifically, the emotion recognition system analyzes the teacher's facial expressions and tone of voice, and sends the emotion data to a server. The server has the function to slightly adjust the teacher's teaching methods based on this data. In addition, it supports self-emotional management by providing real-time visual feedback on the teacher's emotional state.
[0138] For example, if a student asks a question during class, the device captures the audio, which is quickly analyzed on the server to provide the student with an immediate and appropriate answer. Furthermore, emotion recognition technology is used to monitor the teacher's stress level, and if high stress is detected, an alert is sent to the teacher as a warning. This data is later used to optimize the entire system and also helps maintain the teacher's mental health.
[0139] Thus, this system aims to improve the accuracy of individualized instruction, reduce the burden on teachers, and enhance the quality of education by utilizing multifaceted data, including feedback from students.
[0140] The following describes the processing flow.
[0141] Step 1:
[0142] The device collects students' voices in the classroom using a microphone and simultaneously captures video data such as facial expressions and movements using a camera. This makes it possible to acquire audio and video information in real time.
[0143] Step 2:
[0144] The server receives the audio data sent from the device and converts it into text data through a speech recognition engine. This text conversion allows for a more detailed understanding of the student's questions.
[0145] Step 3:
[0146] The server analyzes the converted text data using natural language processing to identify the intent and purpose of the questions. Based on these analysis results, appropriate answers and additional guidance for students are formulated.
[0147] Step 4:
[0148] The server receives video data transmitted from the terminal and analyzes it using deep learning. Here, the emotional state and concentration level of students are evaluated from their facial expressions and gestures, allowing for real-time monitoring of their learning progress.
[0149] Step 5:
[0150] Based on the analysis results, the server uses content generation methods to create personalized learning content for each student. Optimized content is provided based on historical learning data and current understanding.
[0151] Step 6:
[0152] The system uses emotion recognition to obtain user (teacher) emotion data, which is then sent to a server for analysis to understand the user's emotional state. The server generates feedback to adjust teaching content and methods according to stress levels and emotional changes.
[0153] Step 7:
[0154] Based on feedback from the server, users can monitor their own emotions and the students' learning progress to create effective teaching plans. Real-time visualization of emotional data supports emotional management during instruction.
[0155] Step 8:
[0156] All feedback data is stored on the server and continuously used to optimize the entire system. This ensures continuous improvement in the quality and efficiency of education.
[0157] (Example 2)
[0158] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0159] In educational settings, individualized support tailored to each student's learning situation and emotional state is required, but in reality, this is difficult to achieve due to the heavy burden on teachers. Furthermore, teachers' well-being can negatively impact the quality of instruction. Additionally, prompt and accurate responses to student questions are necessary. A system that comprehensively addresses these challenges is needed.
[0160] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0161] In this invention, the server includes recognition means for converting student voice information into text data, processing means for analyzing the content of student questions, and analysis means for analyzing student video information and evaluating learning progress. This enables learning support tailored to each student's learning situation, and also improves the quality of instruction through analysis of the teacher's emotional state and feedback function. Furthermore, by providing a function to respond to student questions immediately, it supports the rapid progress of lessons.
[0162] "Acquisition means" refers to a device or function that acquires students' audio and video information in real time.
[0163] "Recognition means" refers to software or a device that converts acquired audio information into text data.
[0164] "Processing means" refers to the functions and processes for analyzing the content of students' questions from text data.
[0165] "Video acquisition means" refers to a device or function for acquiring video information of students.
[0166] The "analysis method" is a function that evaluates students' learning progress and emotional state based on acquired video information.
[0167] "Generation means" refers to a function that generates individual learning information based on the learning history and analysis results.
[0168] "Means of provision" refers to a function or device for providing students with generated learning information.
[0169] A "learning tool" is a function that accumulates feedback information and new learning information to optimize the entire system.
[0170] A "means of response" refers to a function or process for providing an immediate and appropriate answer to a student's question.
[0171] "Analysis means" refers to functions and processes for analyzing the user's emotional state and adjusting the teaching method accordingly.
[0172] "Means of providing feedback" refers to a function that notifies users of analysis results and emotional states in real time.
[0173] This system is designed to provide efficient learning support in educational settings. Its main components include terminals, servers, and users (primarily teachers).
[0174] The terminal is installed in the classroom and is a device for acquiring students' audio and video in real time. It incorporates an audio acquisition device and a camera to reliably capture students' audio and video information. This data is transmitted to a server via the internet or a local network.
[0175] The server plays a central role in data processing. Audio data is converted into text data using speech recognition software (e.g., a general-purpose speech recognition API). Next, natural language processing is performed using a generative AI model to analyze the content of the students' questions. Based on the analysis results, a prompt such as "Please create an explanation for the student's question" is input to the generative AI model to obtain an appropriate response.
[0176] Meanwhile, the video data is analyzed by emotion recognition software utilizing deep learning technology (e.g., open-source emotion recognition libraries). This allows for the evaluation of students' emotional states and concentration levels from their facial expressions and movements. Furthermore, the server generates personalized learning content based on the students' learning history and evaluation results, and provides it to the students via their devices.
[0177] The user, acting as a teacher, can analyze their own emotional state using an emotion engine. The device recognizes the teacher's voice and facial expressions, and sends the results to the server, which then adjusts the teaching method. Furthermore, the server has a function to provide feedback on the teacher's emotional state, supporting improved teaching quality and stress management.
[0178] For example, if a student asks a question such as, "Please explain photosynthesis in plants," the device captures the audio and analyzes it on a server. Based on the analysis, a generating AI model provides a specific and clear explanation, which is then immediately answered for the student. In addition, if a teacher's stress level is high, a warning alert is issued, and appropriate measures are taken depending on the cause.
[0179] In this way, the entire system supports individual student learning and effectively assists teachers in their instruction, thereby contributing to an improvement in the quality of education.
[0180] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0181] Step 1:
[0182] The device acquires audio and video of students in the classroom. The input consists of student speech and actions. Audio is collected via a microphone device, and video is captured by a camera. This data is then appropriately packaged for transmission to the server.
[0183] Step 2:
[0184] The server receives audio data transmitted from the terminal. The input here is the student's voice signal. The server uses speech recognition software to convert this audio signal into text data. The output is text representing what the student said.
[0185] Step 3:
[0186] The server analyzes the converted text data. A generative AI model is used for this, with the converted text as input. It utilizes natural language processing techniques to understand the meaning of the text and identifies the student's question as output. Specifically, it uses prompts to instruct the generative AI model to generate answers appropriate to the question.
[0187] Step 4:
[0188] The server generates answers to student questions. The input consists of the previously identified question and associated prompt. A generative AI model is used to generate appropriate answers, obtaining the answer text as output. This answer is then ready to be presented to the student.
[0189] Step 5:
[0190] The device provides feedback to the student based on the answers provided by the server. The input is the answer data sent from the server. The device conveys this data to the student as audio or text display.
[0191] Step 6:
[0192] The server analyzes video data of students sent from their terminals. In this step, deep learning is used to analyze students' facial expressions and movements. The input is video data, and the output generates evaluation results of the students' emotional state and concentration level.
[0193] Step 7:
[0194] The server records students' learning progress based on evaluation results and generates individualized learning content as needed. The input consists of emotional state and learning history data, and the output generated based on this data is personalized learning content. This enables more effective learning support for students.
[0195] Step 8:
[0196] The terminal acquires the emotional state of the user (teacher) and sends it to the server. The input here is the teacher's voice and facial expression information. The server uses an emotion engine to analyze this data and identifies the teacher's emotional state as output.
[0197] Step 9:
[0198] The server provides feedback tailored to the teacher's emotional state. The input is analyzed emotional data, and the output includes suggestions for adjusting teaching methods for the teacher. In this step, the teacher receives real-time feedback on their own emotional state and is prompted to take necessary actions.
[0199] (Application Example 2)
[0200] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0201] The problem that this invention aims to solve is to provide an advanced and flexible system for individually supporting students' learning in educational settings. In particular, there is a need to create a more effective educational environment by optimizing learning content while taking into account students' emotional states and levels of concentration, and by monitoring teachers' emotional states.
[0202] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0203] In this invention, the server includes a feedback means for collecting student responses and providing emotional feedback to teachers, an analysis means for analyzing the emotional state of educators and supporting improvement of instruction, and a means for analyzing emotions from video data and providing optimized educational support. This enables learning support tailored to each student, improvement of teachers' instructional skills, and effective stress management.
[0204] "Means of acquisition" refers to devices or methods for collecting students' audio or video information.
[0205] "Recognition means" refers to devices or methods that have the function of analyzing acquired audio information and converting it into text data.
[0206] "Processing means" refers to devices or methods for analyzing text data to identify the content of students' questions and derive appropriate responses.
[0207] "Analysis means" refers to devices or methods for analyzing video information and evaluating students' learning progress and emotional state.
[0208] "Generation means" refers to a device or method for creating individualized learning content based on students' learning history and analysis results.
[0209] "Means of delivery" refers to the devices and methods used to deliver the generated learning content to students.
[0210] A "feedback tool" is a device or method that collects student responses and provides emotional feedback to the teacher.
[0211] A "learning tool" is a device or method for accumulating feedback data and new learning data to optimize the entire system.
[0212] "Detection means" refers to devices or methods for identifying students' poor health or abnormal behavior and generating alerts in advance.
[0213] A "response mechanism" refers to a device or method for providing immediate answers to students' questions.
[0214] This invention provides student learning support through a system using an AI robot teacher for educational purposes. The system mainly consists of a server, terminals, and various analysis means.
[0215] The server receives audio and video information from the student's device and processes this information. The audio information is converted into text data using the Google Cloud Speech-to-Text API, and the OpenAI GPT model is used for natural language processing. At this stage, the server analyzes the student's question and generates an appropriate answer.
[0216] The server uses Microsoft® Azure® Emotion API to perform emotion analysis on video information. This allows for the evaluation of students' learning progress and emotional states, and the provision of personalized learning content and support.
[0217] Furthermore, the feedback system notifies teachers of their emotional state in real time, supporting the improvement of their teaching skills. The feedback function includes an interface that visually shows how the teacher's own emotions are affecting their lessons.
[0218] For example, if a student asks, "I don't know how to find the area of this triangle," the system recognizes the voice and provides a real-time answer such as, "The area of a triangle can be calculated by base × height ÷ 2. For example, if the base is 5 cm and the height is 4 cm, the area is 10 square centimeters." Furthermore, support is provided based on the emotional state of the student and teacher.
[0219] Examples of prompts to input into a generative AI model:
[0220] A student asked, "How do you find the area of a triangle?" Please provide an educational and easy-to-understand explanation to answer this question.
[0221] This allows the entire system to provide more effective educational support based on two-way feedback between students and teachers.
[0222] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0223] Step 1:
[0224] The device acquires student audio and video information. The input is the student's raw audio and video. The device uses its built-in camera and microphone to collect this data in real time and send it to the server. The output is the transmission of data to the server.
[0225] Step 2:
[0226] The server converts the received audio information into text data using the Google Cloud Speech-to-Text API. The input is the student's audio data. The server uses speech recognition technology to analyze the audio data and generate text data. The output is the student's question in text form.
[0227] Step 3:
[0228] The server inputs text data into an OpenAI GPT model, analyzes the question, and generates an answer. The input consists of student questions transcribed from speech. Natural language processing techniques are used to process the data and generate appropriate answers. The output is the answer text that can be provided to the student.
[0229] Step 4:
[0230] The server simultaneously analyzes the received video data using the Microsoft Azure Emotion API to evaluate the students' emotional state. The input is the students' video data. Deep learning technology is used to recognize emotions and determine the students' emotional state and level of concentration. The output is data indicating the students' emotional state.
[0231] Step 5:
[0232] The server creates a plan to provide individualized learning content and support based on the generated responses and emotional state data. The input is the generated responses and emotional state data. The server then develops an individualized learning plan by comparing it with the student's learning history and other factors. The output is the learning content provided to the student.
[0233] Step 6:
[0234] The terminal delivers the provided learning content to the students. The input is learning content data sent from the server. The terminal displays the content to the students via its display and provides learning support. The output is learning information displayed to the students.
[0235] Step 7:
[0236] The server provides visual feedback to the user (teacher) regarding the student's learning progress and the teacher's own emotional state using feedback mechanisms. The input consists of student learning progress and teacher's emotional data. Based on this data, the server generates feedback information for the teacher. The output is the feedback information provided to the teacher.
[0237] Through these steps, the server, terminal, and user work together to provide detailed, personalized learning support for each student.
[0238] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0239] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0240] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0241] [Second Embodiment]
[0242] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0243] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0244] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0245] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0246] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0247] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0248] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0249] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0250] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0251] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0252] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0253] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0254] This invention is a learning support system using an AI robot teacher in educational settings, aiming to reduce the burden on teachers and improve the quality of individualized instruction for students. This system mainly consists of terminals and a server.
[0255] The terminals are installed in the classrooms and are equipped with microphones to capture student voices in real time, and cameras to monitor students' facial expressions and behavior. As a result, audio and video information is acquired simultaneously and transmitted to the server.
[0256] The server uses a speech recognition engine to convert the audio information sent from the terminal into text data. At this stage, questions and requests are extracted from the audio. Furthermore, natural language processing is used to analyze the text data and understand the intent behind the student's questions. This analysis generates appropriate answers and learning content.
[0257] The server uses deep learning algorithms to analyze the video data, analyzing students' emotional states and concentration levels. This allows teachers to identify students who require individual attention and to evaluate the overall learning environment of the class. Furthermore, based on this data, an anomaly detection system immediately generates an alert if it detects any student distress or abnormal behavior, notifying the teacher.
[0258] Furthermore, the server uses a generative AI model to generate personalized learning content based on students' past learning history and progress data. This content is adjusted in difficulty and content according to the student's level of understanding, providing optimal learning materials tailored to individual needs.
[0259] Teachers, as users, can utilize the data and reports provided by their devices to understand their students' situations and develop effective teaching plans. This frees teachers from the detailed responses required in daily lessons, allowing them to dedicate more time and effort to individualized instruction.
[0260] Thus, the AI robot teacher system of the present invention can contribute to improving the efficiency and quality of educational settings by utilizing speech recognition and video analysis to automate student learning support.
[0261] The following describes the processing flow.
[0262] Step 1:
[0263] The device uses a microphone to capture audio information emitted by students in the classroom, while simultaneously capturing students' facial expressions and movements with a camera. This allows for the collection of audio and video data in real time.
[0264] Step 2:
[0265] The device sends the collected audio data to the server. The server receives this data and converts it into text data using a speech recognition engine. The recognized text is then processed to facilitate smooth responses to the teacher's questions.
[0266] Step 3:
[0267] The server analyzes the converted text data using natural language processing techniques to understand the intent and content of the questions. Based on this analysis, it generates answers to help students understand the questions.
[0268] Step 4:
[0269] The terminal sends video data to the server for analysis. The server uses deep learning technology to analyze the students' facial expressions and movements to evaluate their emotional state and level of concentration.
[0270] Step 5:
[0271] The server uses anomaly detection methods based on the analysis results to detect student malaise or abnormal behavior. If detected, an alert is immediately generated and notified to the user.
[0272] Step 6:
[0273] The server references the student's past learning history and generates optimized, personalized learning content using a generative AI model. The generated content is customized to the student's current level of understanding and progress.
[0274] Step 7:
[0275] The device provides students with generated learning content and receives real-time feedback from students through the device. This feedback is used to generate the next learning content.
[0276] Step 8:
[0277] Users can use real-time data and reports provided by their devices to understand students' learning progress and plan effective instruction. Teachers can more easily provide instruction based on quantitative data.
[0278] (Example 1)
[0279] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0280] There is a need to reduce the burden on teachers in educational settings and improve the quality of individualized instruction for students. Traditional methods have made it difficult to effectively respond to students' learning situations and individual needs, and it has also been difficult to grasp the condition of individual students in large groups. Furthermore, there are insufficient means to detect and respond to student distress or abnormal behavior at an early stage, so these challenges need to be addressed.
[0281] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0282] In this invention, the server includes a recording device for acquiring audio, a conversion device for converting audio information into text, and an analysis device for analyzing text information and extracting intent. This enables effective individualized instruction tailored to the student's learning situation and needs.
[0283] A "recording device" is a device used to collect audio information from students.
[0284] The "conversion device" is a device for converting the acquired voice information into character data.
[0285] The "analysis device" is a device for examining character data in detail and extracting the intentions and question contents of students.
[0286] The "imaging device" is a device for acquiring video information of students.
[0287] The "judgment device" is a device for analyzing video information of students and evaluating the learning situation and emotional state.
[0288] The "generation device" is a device for creating individual teaching materials based on the learning history and analysis results of students.
[0289] The "providing device" is a device for supplying the generated teaching materials to students.
[0290] [[ID=2I3]] The "storage device" is a device for storing learning data and optimizing the system.
[0291] The "detection device" is a device for detecting the discomfort and abnormal behavior of students and generating an alert in advance.
[0292] The "response device" is a device for immediately providing an answer to a student's question.
[0293] The present invention is an AI system aimed at supporting classes in the educational field. It mainly consists of a terminal installed in the classroom and a server for data processing. The terminal includes a high-sensitivity microphone for acquiring voice information of students and a high-resolution camera for recording video information of students.
[0294] The server uses a speech recognition engine to convert the audio information sent from the terminal into text data. A common example of software used in this speech recognition process is a cloud-based speech recognition service. The text data is then analyzed using natural language processing techniques to clarify the intent behind the student's questions and requests. For example, a natural language processing library might be used for this process.
[0295] Regarding video data, the server utilizes deep learning to evaluate students' emotional states and concentration levels. The algorithms used in this evaluation process include open-source machine learning frameworks. Furthermore, if any distress or abnormal behavior is detected based on the analysis results, an alert is immediately generated and notified to the teacher's terminal.
[0296] Furthermore, the server analyzes students' past learning history and uses a generative AI model to generate personalized learning materials. These materials are adjusted according to the student's level of understanding, providing a learning experience tailored to individual learning needs.
[0297] For example, if a student says, "I don't understand the meaning of differentiation," the system can analyze their intent, generate and provide learning materials that include the basic concepts of differentiation. An example of a prompt from this system is: "When a student asks a question about something they don't understand during class, analyze their intent and generate appropriate learning content."
[0298] As described above, the present invention reduces the burden on teachers and makes it possible to efficiently provide individualized instruction tailored to each student.
[0299] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0300] Step 1:
[0301] The terminal uses a microphone to acquire the voices of students in the classroom. As input, the voices of students reach the terminal. As a specific operation, voice data is collected when a student asks "How do I solve this equation?" during the class. As output, the acquired voice data is sent to the server.
[0302] Step 2:
[0303] The server processes the transmitted voice data with a voice recognition engine and converts it into character data. As input, the voice data received from the terminal is referenced. As a specific operation, the voice recognition engine recognizes the voice "How do I solve this equation?" and generates "How do I solve this equation?" as the corresponding text. As output, the generated text data is obtained.
[0304] Step 3:
[0305] The server analyzes the generated text data with natural language processing means and extracts the intention of the student's question. As input, the text data obtained from voice recognition is used. As a specific operation, by analyzing with natural language processing technology, the intention of "wanting to know the solution method of the equation" is identified. As output, the intention of the question becomes clear.
[0306] Step 4:
[0307] The server processes the video data transmitted from the terminal with a deep learning algorithm for sentiment analysis and concentration evaluation. As input, the video data received from the terminal is used. As a specific operation, the video is analyzed, and the emotional state of "being confused" is detected from the expressions and postures of the students. As output, the analyzed emotional state information is obtained.
[0308] Step 5:
[0309] The server uses a generative AI model to generate learning content tailored to each student, based on the intent and emotional state of the question. The input includes the student's past learning history and analyzed data. Specifically, the AI model automatically creates materials including "basic equations and how to solve them." The output is personalized learning content.
[0310] Step 6:
[0311] The user, the teacher, reviews the generated learning content and analysis results, and provides feedback to students based on them. The inputs include content provided by the server and student status information. Specifically, the teacher utilizes the information from the system to provide individualized instruction to students during class. The output is the provision of appropriate guidance and support to students.
[0312] (Application Example 1)
[0313] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0314] To improve factory production efficiency, it is necessary to accurately understand workers' actions and emotional states and provide them with the most appropriate work procedures in a timely manner. However, conventional systems have difficulty analyzing workers' actions and emotions in real time, and are unable to detect potential hazards in advance and respond appropriately, resulting in insufficient improvements in work efficiency and safety.
[0315] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0316] In this invention, the server includes voice acquisition means for acquiring worker movement information, video analysis means for analyzing worker video information and evaluating production efficiency, and anomaly detection means for alerting inappropriate movements in advance. This makes it possible to grasp the worker's movements and emotions in real time, provide optimal work procedures, and avoid potential hazards.
[0317] "Workers" refer to people who assemble parts or process products within a factory, and their actions and emotional states are the targets for management and optimization.
[0318] "Motion information" refers to data about the movements and work attitude of workers, and is real-time information acquired through sensors such as cameras.
[0319] A "voice acquisition means" is a device that has the function of collecting the voice emitted by a worker as digital data using a microphone or other device.
[0320] "Speech recognition means" refers to a technology that has the function of converting acquired speech data into text data and analyzes the content of what the worker says as text information.
[0321] "Natural language processing means" refers to technologies that analyze text data generated by speech recognition means and understand its context and intent.
[0322] "Video acquisition means" refers to a device that has the function of recording the actions of workers as video using a camera or similar device, and is used for analyzing their movements and emotions.
[0323] "Video analysis means" refers to technology that processes acquired video data to evaluate the productivity and emotional state of workers.
[0324] A "content generation method" is a system that has the function of generating individually optimized work content based on the results of analysis of the worker's actions and emotions.
[0325] A "content delivery means" is a device that has the function of communicating the generated optimization work content to workers and supporting them in carrying out their tasks.
[0326] "Learning methods" refer to technologies that accumulate feedback data and newly acquired work data to improve the accuracy and efficiency of the system.
[0327] An "anomaly detection method" is a technology that detects inappropriate actions or potential hazards by workers in real time and issues an immediate alert.
[0328] A "real-time response system" is a technology that provides optimal work procedures and advice in response to the worker's instructions and situation.
[0329] To implement this invention, a system is configured in which a server, terminal, and user cooperate to evaluate and optimize the worker's actions in real time. The server continuously acquires data from the work site by utilizing voice acquisition means and video acquisition means to collect worker action information. The voice acquisition means uses a microphone to capture the worker's speech and converts it into text data using speech recognition means. Natural language processing software analyzes the worker's intentions and instructions from the text data.
[0330] Simultaneously, the video data collected by the video acquisition system is processed using video analysis tools that apply deep learning algorithms to evaluate the workers' actions and emotions. This makes it possible to grasp production efficiency and the workers' psychological state in real time.
[0331] Based on the analysis results, the server uses a generation AI model to create individually optimized work instructions in the content generation means, and provides appropriate work instructions to the user (worker) via the content provision means. In addition, if the anomaly detection means detects inappropriate operation or potential danger, an alert is immediately generated and notified to the worker and supervisor.
[0332] The learning method analyzes accumulated feedback data and newly acquired work data to continuously optimize the system. The hardware used in this process includes professional cameras, high-sensitivity microphones, and a high-performance computing platform located on a server as computing resources. Software such as OpenCV, Keras, and SpeechRecognition are employed.
[0333] As a concrete example, when a worker is assembling parts, the server detects the worker's anxious expression using video analysis, and a generative AI model generates encouraging comments and provides voice feedback through a content delivery system, such as "Please feel free to contact us if you have any problems." An example of a prompt message would be, "Please propose a new strategy to optimize the next automation process." In this way, the system plays a role in enhancing the smoothness and safety of work.
[0334] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0335] Step 1:
[0336] The server receives voice data from the worker, acquired via microphone, as input from the terminal. This voice data is then used as input for speech recognition, digital signal processing is performed, and it is converted into text data for output. Specifically, the characteristics of the voice waveform are analyzed, and recognition is performed using a language model.
[0337] Step 2:
[0338] The server uses the text data output in step 1 as input and processes it with natural language processing. Here, text analysis is performed to identify the intent behind the worker's instructions and questions, and data is output to determine actions based on that intent. Context and intent are interpreted using natural language processing technology.
[0339] Step 3:
[0340] The server acquires video information of workers as input via cameras connected to terminals. This data is then processed by a video analysis system, and a deep learning algorithm is used to evaluate production efficiency and the emotional state of the workers, outputting the evaluation results. Facial expressions and movement patterns are analyzed from the video data to infer psychological state and work speed.
[0341] Step 4:
[0342] The server uses the intent and evaluation results obtained in steps 2 and 3 as input to generate individually optimized work content using a generative AI model. This generated work content is output by a content generation means. At this stage, prompts are used to generate work procedures and advice as specific instructions.
[0343] Step 5:
[0344] The server provides the content generated in step 4 to the user (worker) via the content delivery means. The outputted work instructions and advice are transmitted to the worker as audio or video, supporting the worker in properly performing their tasks.
[0345] Step 6:
[0346] The server stores feedback data collected during the work process and newly acquired work data as input to its learning mechanism. This allows for long-term data analysis to optimize the system and improve its accuracy. Statistical analysis of the data and machine learning are used to improve the accuracy of future work instructions.
[0347] Step 7:
[0348] The server uses anomaly detection mechanisms to detect inappropriate worker behavior and potential hazards in real time. Based on this input, it immediately generates and outputs alerts, notifying the user (worker) and relevant supervisors. This ensures workplace safety and enables a rapid response.
[0349] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0350] This invention aims to achieve more effective learning support by combining an emotion engine with a system that utilizes AI robot teachers in educational settings. This system mainly consists of a terminal, a server, and emotion recognition means.
[0351] The terminal has the ability to acquire students' audio and video in real time within the classroom and transmits this data to a server. The server converts the received audio data into text data using speech recognition technology and analyzes the content of the questions using natural language processing technology. Based on these results, specific answers and individualized learning content are generated to support students' learning.
[0352] The server analyzes the video data using deep learning technology to evaluate students' emotional states and concentration levels based on their facial expressions and movements. This evaluation allows the system to track students' learning progress, and individualized support is provided as needed through content generation methods.
[0353] The emotion engine is used specifically to analyze the emotional state of the user, the teacher. This allows for feedback that helps maintain or improve the quality of instruction based on the teacher's own emotions.
[0354] Specifically, the emotion recognition system analyzes the teacher's facial expressions and tone of voice, and sends the emotion data to a server. The server has the function to slightly adjust the teacher's teaching methods based on this data. In addition, it supports self-emotional management by providing real-time visual feedback on the teacher's emotional state.
[0355] For example, if a student asks a question during class, the device captures the audio, which is quickly analyzed on the server to provide the student with an immediate and appropriate answer. Furthermore, emotion recognition technology is used to monitor the teacher's stress level, and if high stress is detected, an alert is sent to the teacher as a warning. This data is later used to optimize the entire system and also helps maintain the teacher's mental health.
[0356] Thus, this system aims to improve the accuracy of individualized instruction, reduce the burden on teachers, and enhance the quality of education by utilizing multifaceted data, including feedback from students.
[0357] The following describes the processing flow.
[0358] Step 1:
[0359] The device collects students' voices in the classroom using a microphone and simultaneously captures video data such as facial expressions and movements using a camera. This makes it possible to acquire audio and video information in real time.
[0360] Step 2:
[0361] The server receives the audio data sent from the device and converts it into text data through a speech recognition engine. This text conversion allows for a more detailed understanding of the student's questions.
[0362] Step 3:
[0363] The server analyzes the converted text data using natural language processing to identify the intent and purpose of the questions. Based on these analysis results, appropriate answers and additional guidance for students are formulated.
[0364] Step 4:
[0365] The server receives video data transmitted from the terminal and analyzes it using deep learning. Here, the emotional state and concentration level of students are evaluated from their facial expressions and gestures, allowing for real-time monitoring of their learning progress.
[0366] Step 5:
[0367] Based on the analysis results, the server uses content generation methods to create personalized learning content for each student. Optimized content is provided based on historical learning data and current understanding.
[0368] Step 6:
[0369] The system uses emotion recognition to obtain user (teacher) emotion data, which is then sent to a server for analysis to understand the user's emotional state. The server generates feedback to adjust teaching content and methods according to stress levels and emotional changes.
[0370] Step 7:
[0371] Based on feedback from the server, users can monitor their own emotions and the students' learning progress to create effective teaching plans. Real-time visualization of emotional data supports emotional management during instruction.
[0372] Step 8:
[0373] All feedback data is stored on the server and continuously used to optimize the entire system. This ensures continuous improvement in the quality and efficiency of education.
[0374] (Example 2)
[0375] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0376] In educational settings, individualized support tailored to each student's learning situation and emotional state is required, but in reality, this is difficult to achieve due to the heavy burden on teachers. Furthermore, teachers' well-being can negatively impact the quality of instruction. Additionally, prompt and accurate responses to student questions are necessary. A system that comprehensively addresses these challenges is needed.
[0377] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0378] In this invention, the server includes recognition means for converting student voice information into text data, processing means for analyzing the content of student questions, and analysis means for analyzing student video information and evaluating learning progress. This enables learning support tailored to each student's learning situation, and also improves the quality of instruction through analysis of the teacher's emotional state and feedback function. Furthermore, by providing a function to respond to student questions immediately, it supports the rapid progress of lessons.
[0379] "Acquisition means" refers to a device or function that acquires students' audio and video information in real time.
[0380] "Recognition means" refers to software or a device that converts acquired audio information into text data.
[0381] "Processing means" refers to the functions and processes for analyzing the content of students' questions from text data.
[0382] "Video acquisition means" refers to a device or function for acquiring video information of students.
[0383] The "analysis method" is a function that evaluates students' learning progress and emotional state based on acquired video information.
[0384] "Generation means" refers to a function that generates individual learning information based on the learning history and analysis results.
[0385] "Means of provision" refers to a function or device for providing students with generated learning information.
[0386] A "learning tool" is a function that accumulates feedback information and new learning information to optimize the entire system.
[0387] A "means of response" refers to a function or process for providing an immediate and appropriate answer to a student's question.
[0388] "Analysis means" refers to functions and processes for analyzing the user's emotional state and adjusting the teaching method accordingly.
[0389] "Means of providing feedback" refers to a function that notifies users of analysis results and emotional states in real time.
[0390] This system is designed to provide efficient learning support in educational settings. Its main components include terminals, servers, and users (primarily teachers).
[0391] The terminal is installed in the classroom and is a device for acquiring students' audio and video in real time. It incorporates an audio acquisition device and a camera to reliably capture students' audio and video information. This data is transmitted to a server via the internet or a local network.
[0392] The server plays a central role in data processing. Audio data is converted into text data using speech recognition software (e.g., a general-purpose speech recognition API). Next, natural language processing is performed using a generative AI model to analyze the content of the students' questions. Based on the analysis results, a prompt such as "Please create an explanation for the student's question" is input to the generative AI model to obtain an appropriate response.
[0393] Meanwhile, the video data is analyzed by emotion recognition software utilizing deep learning technology (e.g., open-source emotion recognition libraries). This allows for the evaluation of students' emotional states and concentration levels from their facial expressions and movements. Furthermore, the server generates personalized learning content based on the students' learning history and evaluation results, and provides it to the students via their devices.
[0394] The user, acting as a teacher, can analyze their own emotional state using an emotion engine. The device recognizes the teacher's voice and facial expressions, and sends the results to the server, which then adjusts the teaching method. Furthermore, the server has a function to provide feedback on the teacher's emotional state, supporting improved teaching quality and stress management.
[0395] For example, if a student asks a question such as, "Please explain photosynthesis in plants," the device captures the audio and analyzes it on a server. Based on the analysis, a generating AI model provides a specific and clear explanation, which is then immediately answered for the student. In addition, if a teacher's stress level is high, a warning alert is issued, and appropriate measures are taken depending on the cause.
[0396] In this way, the entire system supports individual student learning and effectively assists teachers in their instruction, thereby contributing to an improvement in the quality of education.
[0397] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0398] Step 1:
[0399] The device acquires audio and video of students in the classroom. The input consists of student speech and actions. Audio is collected via a microphone device, and video is captured by a camera. This data is then appropriately packaged for transmission to the server.
[0400] Step 2:
[0401] The server receives audio data transmitted from the terminal. The input here is the student's voice signal. The server uses speech recognition software to convert this audio signal into text data. The output is text representing what the student said.
[0402] Step 3:
[0403] The server analyzes the converted text data. A generative AI model is used for this, with the converted text as input. It utilizes natural language processing techniques to understand the meaning of the text and identifies the student's question as output. Specifically, it uses prompts to instruct the generative AI model to generate answers appropriate to the question.
[0404] Step 4:
[0405] The server generates answers to student questions. The input consists of the previously identified question and associated prompt. A generative AI model is used to generate appropriate answers, obtaining the answer text as output. This answer is then ready to be presented to the student.
[0406] Step 5:
[0407] The device provides feedback to the student based on the answers provided by the server. The input is the answer data sent from the server. The device conveys this data to the student as audio or text display.
[0408] Step 6:
[0409] The server analyzes video data of students sent from their terminals. In this step, deep learning is used to analyze students' facial expressions and movements. The input is video data, and the output generates evaluation results of the students' emotional state and concentration level.
[0410] Step 7:
[0411] The server records students' learning progress based on evaluation results and generates individualized learning content as needed. The input consists of emotional state and learning history data, and the output generated based on this data is personalized learning content. This enables more effective learning support for students.
[0412] Step 8:
[0413] The terminal acquires the emotional state of the user (teacher) and sends it to the server. The input here is the teacher's voice and facial expression information. The server uses an emotion engine to analyze this data and identifies the teacher's emotional state as output.
[0414] Step 9:
[0415] The server provides feedback tailored to the teacher's emotional state. The input is analyzed emotional data, and the output includes suggestions for adjusting teaching methods for the teacher. In this step, the teacher receives real-time feedback on their own emotional state and is prompted to take necessary actions.
[0416] (Application Example 2)
[0417] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0418] The problem that this invention aims to solve is to provide an advanced and flexible system for individually supporting students' learning in educational settings. In particular, there is a need to create a more effective educational environment by optimizing learning content while taking into account students' emotional states and levels of concentration, and by monitoring teachers' emotional states.
[0419] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0420] In this invention, the server includes a feedback means for collecting student responses and providing emotional feedback to teachers, an analysis means for analyzing the emotional state of educators and supporting improvement of instruction, and a means for analyzing emotions from video data and providing optimized educational support. This enables learning support tailored to each student, improvement of teachers' instructional skills, and effective stress management.
[0421] "Means of acquisition" refers to devices or methods for collecting students' audio or video information.
[0422] "Recognition means" refers to devices or methods that have the function of analyzing acquired audio information and converting it into text data.
[0423] "Processing means" refers to devices or methods for analyzing text data to identify the content of students' questions and derive appropriate responses.
[0424] "Analysis means" refers to devices or methods for analyzing video information and evaluating students' learning progress and emotional state.
[0425] "Generation means" refers to a device or method for creating individualized learning content based on students' learning history and analysis results.
[0426] "Means of delivery" refers to the devices and methods used to deliver the generated learning content to students.
[0427] A "feedback tool" is a device or method that collects student responses and provides emotional feedback to the teacher.
[0428] A "learning tool" is a device or method for accumulating feedback data and new learning data to optimize the entire system.
[0429] "Detection means" refers to devices or methods for identifying students' poor health or abnormal behavior and generating alerts in advance.
[0430] A "response mechanism" refers to a device or method for providing immediate answers to students' questions.
[0431] This invention provides student learning support through a system using an AI robot teacher for educational purposes. The system mainly consists of a server, terminals, and various analysis means.
[0432] The server receives audio and video information from the student's device and processes this information. The audio information is converted into text data using the Google Cloud Speech-to-Text API, and OpenAI's GPT model is used for natural language processing. At this stage, the server analyzes the student's question and generates an appropriate answer.
[0433] The server uses the Microsoft Azure Emotion API to perform emotion analysis on the video data. This allows the server to evaluate students' learning progress and emotional state, and then provide personalized learning content and support.
[0434] Furthermore, the feedback system notifies teachers of their emotional state in real time, supporting the improvement of their teaching skills. The feedback function includes an interface that visually shows how the teacher's own emotions are affecting their lessons.
[0435] For example, if a student asks, "I don't know how to find the area of this triangle," the system recognizes the voice and provides a real-time answer such as, "The area of a triangle can be calculated by base × height ÷ 2. For example, if the base is 5 cm and the height is 4 cm, the area is 10 square centimeters." Furthermore, support is provided based on the emotional state of the student and teacher.
[0436] Examples of prompts to input into a generative AI model:
[0437] A student asked, "How do you find the area of a triangle?" Please provide an educational and easy-to-understand explanation to answer this question.
[0438] This allows the entire system to provide more effective educational support based on two-way feedback between students and teachers.
[0439] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0440] Step 1:
[0441] The device acquires student audio and video information. The input is the student's raw audio and video. The device uses its built-in camera and microphone to collect this data in real time and send it to the server. The output is the transmission of data to the server.
[0442] Step 2:
[0443] The server converts the received audio information into text data using the Google Cloud Speech-to-Text API. The input is the student's audio data. The server uses speech recognition technology to analyze the audio data and generate text data. The output is the student's question in text form.
[0444] Step 3:
[0445] The server inputs text data into an OpenAI GPT model, analyzes the question, and generates an answer. The input consists of student questions transcribed from speech. Natural language processing techniques are used to process the data and generate appropriate answers. The output is the answer text that can be provided to the student.
[0446] Step 4:
[0447] The server simultaneously analyzes the received video data using the Microsoft Azure Emotion API to evaluate the students' emotional state. The input is the students' video data. Deep learning technology is used to recognize emotions and determine the students' emotional state and level of concentration. The output is data indicating the students' emotional state.
[0448] Step 5:
[0449] The server creates a plan to provide individualized learning content and support based on the generated responses and emotional state data. The input is the generated responses and emotional state data. The server then develops an individualized learning plan by comparing it with the student's learning history and other factors. The output is the learning content provided to the student.
[0450] Step 6:
[0451] The terminal delivers the provided learning content to the students. The input is learning content data sent from the server. The terminal displays the content to the students via its display and provides learning support. The output is learning information displayed to the students.
[0452] Step 7:
[0453] The server provides visual feedback to the user (teacher) regarding the student's learning progress and the teacher's own emotional state using feedback mechanisms. The input consists of student learning progress and teacher's emotional data. Based on this data, the server generates feedback information for the teacher. The output is the feedback information provided to the teacher.
[0454] Through these steps, the server, terminal, and user work together to provide detailed, personalized learning support for each student.
[0455] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0456] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0457] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0458] [Third Embodiment]
[0459] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0460] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0461] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0462] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0463] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0464] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0465] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0466] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0467] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0468] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0469] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0470] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0471] This invention is a learning support system using an AI robot teacher in educational settings, aiming to reduce the burden on teachers and improve the quality of individualized instruction for students. This system mainly consists of terminals and a server.
[0472] The terminals are installed in the classrooms and are equipped with microphones to capture student voices in real time, and cameras to monitor students' facial expressions and behavior. As a result, audio and video information is acquired simultaneously and transmitted to the server.
[0473] The server uses a speech recognition engine to convert the audio information sent from the terminal into text data. At this stage, questions and requests are extracted from the audio. Furthermore, natural language processing is used to analyze the text data and understand the intent behind the student's questions. This analysis generates appropriate answers and learning content.
[0474] The server uses deep learning algorithms to analyze the video data, analyzing students' emotional states and concentration levels. This allows teachers to identify students who require individual attention and to evaluate the overall learning environment of the class. Furthermore, based on this data, an anomaly detection system immediately generates an alert if it detects any student distress or abnormal behavior, notifying the teacher.
[0475] Furthermore, the server uses a generative AI model to generate personalized learning content based on students' past learning history and progress data. This content is adjusted in difficulty and content according to the student's level of understanding, providing optimal learning materials tailored to individual needs.
[0476] Teachers, as users, can utilize the data and reports provided by their devices to understand their students' situations and develop effective teaching plans. This frees teachers from the detailed responses required in daily lessons, allowing them to dedicate more time and effort to individualized instruction.
[0477] Thus, the AI robot teacher system of the present invention can contribute to improving the efficiency and quality of educational settings by utilizing speech recognition and video analysis to automate student learning support.
[0478] The following describes the processing flow.
[0479] Step 1:
[0480] The device uses a microphone to capture audio information emitted by students in the classroom, while simultaneously capturing students' facial expressions and movements with a camera. This allows for the collection of audio and video data in real time.
[0481] Step 2:
[0482] The device sends the collected audio data to the server. The server receives this data and converts it into text data using a speech recognition engine. The recognized text is then processed to facilitate smooth responses to the teacher's questions.
[0483] Step 3:
[0484] The server analyzes the converted text data using natural language processing techniques to understand the intent and content of the questions. Based on this analysis, it generates answers to help students understand the questions.
[0485] Step 4:
[0486] The terminal sends video data to the server for analysis. The server uses deep learning technology to analyze the students' facial expressions and movements to evaluate their emotional state and level of concentration.
[0487] Step 5:
[0488] The server uses anomaly detection methods based on the analysis results to detect student malaise or abnormal behavior. If detected, an alert is immediately generated and notified to the user.
[0489] Step 6:
[0490] The server references the student's past learning history and generates optimized, personalized learning content using a generative AI model. The generated content is customized to the student's current level of understanding and progress.
[0491] Step 7:
[0492] The device provides students with generated learning content and receives real-time feedback from students through the device. This feedback is used to generate the next learning content.
[0493] Step 8:
[0494] Users can use real-time data and reports provided by their devices to understand students' learning progress and plan effective instruction. Teachers can more easily provide instruction based on quantitative data.
[0495] (Example 1)
[0496] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0497] There is a need to reduce the burden on teachers in educational settings and improve the quality of individualized instruction for students. Traditional methods have made it difficult to effectively respond to students' learning situations and individual needs, and it has also been difficult to grasp the condition of individual students in large groups. Furthermore, there are insufficient means to detect and respond to student distress or abnormal behavior at an early stage, so these challenges need to be addressed.
[0498] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0499] In this invention, the server includes a recording device for acquiring audio, a conversion device for converting audio information into text, and an analysis device for analyzing text information and extracting intent. This enables effective individualized instruction tailored to the student's learning situation and needs.
[0500] A "recording device" is a device used to collect audio information from students.
[0501] A "conversion device" is a device used to convert acquired audio information into text data.
[0502] An "analysis device" is a device that examines text data in detail to extract the student's intentions and the content of their questions.
[0503] A "filming device" is a device used to acquire video information of students.
[0504] A "judgment device" is a device that analyzes video information of students to evaluate their learning progress and emotional state.
[0505] A "generation device" is a device that creates individualized learning materials based on students' learning history and analysis results.
[0506] A "distribution device" is a device used to supply students with the generated teaching materials.
[0507] A "memory device" is a device used to store learning data and optimize the system.
[0508] A "detection device" is a device that detects students' poor health or abnormal behavior and generates alerts in advance.
[0509] A "response device" is a device that provides immediate answers to students' questions.
[0510] This invention is an AI system intended to support classroom instruction in educational settings. It mainly consists of a terminal installed in the classroom and a server for data processing. The terminal includes a high-sensitivity microphone for acquiring student voice information and a high-resolution camera for recording student video information.
[0511] The server uses a speech recognition engine to convert the audio information sent from the terminal into text data. A common example of software used in this speech recognition process is a cloud-based speech recognition service. The text data is then analyzed using natural language processing techniques to clarify the intent behind the student's questions and requests. For example, a natural language processing library might be used for this process.
[0512] Regarding video data, the server utilizes deep learning to evaluate students' emotional states and concentration levels. The algorithms used in this evaluation process include open-source machine learning frameworks. Furthermore, if any distress or abnormal behavior is detected based on the analysis results, an alert is immediately generated and notified to the teacher's terminal.
[0513] Furthermore, the server analyzes students' past learning history and uses a generative AI model to generate personalized learning materials. These materials are adjusted according to the student's level of understanding, providing a learning experience tailored to individual learning needs.
[0514] For example, if a student says, "I don't understand the meaning of differentiation," the system can analyze their intent, generate and provide learning materials that include the basic concepts of differentiation. An example of a prompt from this system is: "When a student asks a question about something they don't understand during class, analyze their intent and generate appropriate learning content."
[0515] As described above, the present invention reduces the burden on teachers and makes it possible to efficiently provide individualized instruction tailored to each student.
[0516] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0517] Step 1:
[0518] The device uses a microphone to capture audio from students in the classroom. The input is the student's voice, which is received by the device. Specifically, during class, asking a question like, "How do you solve this equation?" collects the audio data. The output is the collected audio data, which is sent to the server.
[0519] Step 2:
[0520] The server processes the transmitted audio data using a speech recognition engine and converts it into text data. The audio data received from the terminal is used as input. Specifically, the speech recognition engine recognizes the audio "How do you solve this equation?" and generates the corresponding text "How do you solve this equation?". The generated text data is obtained as output.
[0521] Step 3:
[0522] The server analyzes the generated text data using natural language processing techniques to extract the intent behind the student's question. Text data obtained from speech recognition is used as input. Specifically, the natural language processing technique identifies the intent, "I want to know how to solve this equation." The output reveals the intent behind the question.
[0523] Step 4:
[0524] The server processes video data transmitted from the terminal using a deep learning algorithm for emotion analysis and concentration assessment. The input is the video data received from the terminal. Specifically, it analyzes the video and detects an emotional state such as "confused" from the student's facial expressions and posture. The output is the analyzed emotional state information.
[0525] Step 5:
[0526] The server uses a generative AI model to generate learning content tailored to each student, based on the intent and emotional state of the question. The input includes the student's past learning history and analyzed data. Specifically, the AI model automatically creates materials including "basic equations and how to solve them." The output is personalized learning content.
[0527] Step 6:
[0528] The user, the teacher, reviews the generated learning content and analysis results, and provides feedback to students based on them. The inputs include content provided by the server and student status information. Specifically, the teacher utilizes the information from the system to provide individualized instruction to students during class. The output is the provision of appropriate guidance and support to students.
[0529] (Application Example 1)
[0530] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0531] To improve factory production efficiency, it is necessary to accurately understand workers' actions and emotional states and provide them with the most appropriate work procedures in a timely manner. However, conventional systems have difficulty analyzing workers' actions and emotions in real time, and are unable to detect potential hazards in advance and respond appropriately, resulting in insufficient improvements in work efficiency and safety.
[0532] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0533] In this invention, the server includes voice acquisition means for acquiring worker movement information, video analysis means for analyzing worker video information and evaluating production efficiency, and anomaly detection means for alerting inappropriate movements in advance. This makes it possible to grasp the worker's movements and emotions in real time, provide optimal work procedures, and avoid potential hazards.
[0534] "Workers" refer to people who assemble parts or process products within a factory, and their actions and emotional states are the targets for management and optimization.
[0535] "Motion information" refers to data about the movements and work attitude of workers, and is real-time information acquired through sensors such as cameras.
[0536] A "voice acquisition means" is a device that has the function of collecting the voice emitted by a worker as digital data using a microphone or other device.
[0537] "Speech recognition means" refers to a technology that has the function of converting acquired speech data into text data and analyzes the content of what the worker says as text information.
[0538] "Natural language processing means" refers to technologies that analyze text data generated by speech recognition means and understand its context and intent.
[0539] "Video acquisition means" refers to a device that has the function of recording the actions of workers as video using a camera or similar device, and is used for analyzing their movements and emotions.
[0540] "Video analysis means" refers to technology that processes acquired video data to evaluate the productivity and emotional state of workers.
[0541] A "content generation method" is a system that has the function of generating individually optimized work content based on the results of analysis of the worker's actions and emotions.
[0542] A "content delivery means" is a device that has the function of communicating the generated optimization work content to workers and supporting them in carrying out their tasks.
[0543] "Learning methods" refer to technologies that accumulate feedback data and newly acquired work data to improve the accuracy and efficiency of the system.
[0544] An "anomaly detection method" is a technology that detects inappropriate actions or potential hazards by workers in real time and issues an immediate alert.
[0545] A "real-time response system" is a technology that provides optimal work procedures and advice in response to the worker's instructions and situation.
[0546] To implement this invention, a system is configured in which a server, terminal, and user cooperate to evaluate and optimize the worker's actions in real time. The server continuously acquires data from the work site by utilizing voice acquisition means and video acquisition means to collect worker action information. The voice acquisition means uses a microphone to capture the worker's speech and converts it into text data using speech recognition means. Natural language processing software analyzes the worker's intentions and instructions from the text data.
[0547] Simultaneously, the video data collected by the video acquisition system is processed using video analysis tools that apply deep learning algorithms to evaluate the workers' actions and emotions. This makes it possible to grasp production efficiency and the workers' psychological state in real time.
[0548] Based on the analysis results, the server uses a generation AI model to create individually optimized work instructions in the content generation means, and provides appropriate work instructions to the user (worker) via the content provision means. In addition, if the anomaly detection means detects inappropriate operation or potential danger, an alert is immediately generated and notified to the worker and supervisor.
[0549] The learning method analyzes accumulated feedback data and newly acquired work data to continuously optimize the system. The hardware used in this process includes professional cameras, high-sensitivity microphones, and a high-performance computing platform located on a server as computing resources. Software such as OpenCV, Keras, and SpeechRecognition are employed.
[0550] As a concrete example, when a worker is assembling parts, the server detects the worker's anxious expression using video analysis, and a generative AI model generates encouraging comments and provides voice feedback through a content delivery system, such as "Please feel free to contact us if you have any problems." An example of a prompt message would be, "Please propose a new strategy to optimize the next automation process." In this way, the system plays a role in enhancing the smoothness and safety of work.
[0551] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0552] Step 1:
[0553] The server receives voice data from the worker, acquired via microphone, as input from the terminal. This voice data is then used as input for speech recognition, digital signal processing is performed, and it is converted into text data for output. Specifically, the characteristics of the voice waveform are analyzed, and recognition is performed using a language model.
[0554] Step 2:
[0555] The server uses the text data output in step 1 as input and processes it with natural language processing. Here, text analysis is performed to identify the intent behind the worker's instructions and questions, and data is output to determine actions based on that intent. Context and intent are interpreted using natural language processing technology.
[0556] Step 3:
[0557] The server acquires video information of workers as input via cameras connected to terminals. This data is then processed by a video analysis system, and a deep learning algorithm is used to evaluate production efficiency and the emotional state of the workers, outputting the evaluation results. Facial expressions and movement patterns are analyzed from the video data to infer psychological state and work speed.
[0558] Step 4:
[0559] The server uses the intent and evaluation results obtained in steps 2 and 3 as input to generate individually optimized work content using a generative AI model. This generated work content is output by a content generation means. At this stage, prompts are used to generate work procedures and advice as specific instructions.
[0560] Step 5:
[0561] The server provides the content generated in step 4 to the user (worker) via the content delivery means. The outputted work instructions and advice are transmitted to the worker as audio or video, supporting the worker in properly performing their tasks.
[0562] Step 6:
[0563] The server stores feedback data collected during the work process and newly acquired work data as input to its learning mechanism. This allows for long-term data analysis to optimize the system and improve its accuracy. Statistical analysis of the data and machine learning are used to improve the accuracy of future work instructions.
[0564] Step 7:
[0565] The server uses anomaly detection mechanisms to detect inappropriate worker behavior and potential hazards in real time. Based on this input, it immediately generates and outputs alerts, notifying the user (worker) and relevant supervisors. This ensures workplace safety and enables a rapid response.
[0566] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0567] This invention aims to achieve more effective learning support by combining an emotion engine with a system that utilizes AI robot teachers in educational settings. This system mainly consists of a terminal, a server, and emotion recognition means.
[0568] The terminal has the ability to acquire students' audio and video in real time within the classroom and transmits this data to a server. The server converts the received audio data into text data using speech recognition technology and analyzes the content of the questions using natural language processing technology. Based on these results, specific answers and individualized learning content are generated to support students' learning.
[0569] The server analyzes the video data using deep learning technology to evaluate students' emotional states and concentration levels based on their facial expressions and movements. This evaluation allows the system to track students' learning progress, and individualized support is provided as needed through content generation methods.
[0570] The emotion engine is used specifically to analyze the emotional state of the user, the teacher. This allows for feedback that helps maintain or improve the quality of instruction based on the teacher's own emotions.
[0571] Specifically, the emotion recognition system analyzes the teacher's facial expressions and tone of voice, and sends the emotion data to a server. The server has the function to slightly adjust the teacher's teaching methods based on this data. In addition, it supports self-emotional management by providing real-time visual feedback on the teacher's emotional state.
[0572] For example, if a student asks a question during class, the device captures the audio, which is quickly analyzed on the server to provide the student with an immediate and appropriate answer. Furthermore, emotion recognition technology is used to monitor the teacher's stress level, and if high stress is detected, an alert is sent to the teacher as a warning. This data is later used to optimize the entire system and also helps maintain the teacher's mental health.
[0573] Thus, this system aims to improve the accuracy of individualized instruction, reduce the burden on teachers, and enhance the quality of education by utilizing multifaceted data, including feedback from students.
[0574] The following describes the processing flow.
[0575] Step 1:
[0576] The device collects students' voices in the classroom using a microphone and simultaneously captures video data such as facial expressions and movements using a camera. This makes it possible to acquire audio and video information in real time.
[0577] Step 2:
[0578] The server receives the audio data sent from the device and converts it into text data through a speech recognition engine. This text conversion allows for a more detailed understanding of the student's questions.
[0579] Step 3:
[0580] The server analyzes the converted text data using natural language processing to identify the intent and purpose of the questions. Based on these analysis results, appropriate answers and additional guidance for students are formulated.
[0581] Step 4:
[0582] The server receives video data transmitted from the terminal and analyzes it using deep learning. Here, the emotional state and concentration level of students are evaluated from their facial expressions and gestures, allowing for real-time monitoring of their learning progress.
[0583] Step 5:
[0584] Based on the analysis results, the server uses content generation methods to create personalized learning content for each student. Optimized content is provided based on historical learning data and current understanding.
[0585] Step 6:
[0586] The system uses emotion recognition to obtain user (teacher) emotion data, which is then sent to a server for analysis to understand the user's emotional state. The server generates feedback to adjust teaching content and methods according to stress levels and emotional changes.
[0587] Step 7:
[0588] Based on feedback from the server, users can monitor their own emotions and the students' learning progress to create effective teaching plans. Real-time visualization of emotional data supports emotional management during instruction.
[0589] Step 8:
[0590] All feedback data is stored on the server and continuously used to optimize the entire system. This ensures continuous improvement in the quality and efficiency of education.
[0591] (Example 2)
[0592] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0593] In educational settings, individualized support tailored to each student's learning situation and emotional state is required, but in reality, this is difficult to achieve due to the heavy burden on teachers. Furthermore, teachers' well-being can negatively impact the quality of instruction. Additionally, prompt and accurate responses to student questions are necessary. A system that comprehensively addresses these challenges is needed.
[0594] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0595] In this invention, the server includes recognition means for converting student voice information into text data, processing means for analyzing the content of student questions, and analysis means for analyzing student video information and evaluating learning progress. This enables learning support tailored to each student's learning situation, and also improves the quality of instruction through analysis of the teacher's emotional state and feedback function. Furthermore, by providing a function to respond to student questions immediately, it supports the rapid progress of lessons.
[0596] "Acquisition means" refers to a device or function that acquires students' audio and video information in real time.
[0597] "Recognition means" refers to software or a device that converts acquired audio information into text data.
[0598] "Processing means" refers to the functions and processes for analyzing the content of students' questions from text data.
[0599] "Video acquisition means" refers to a device or function for acquiring video information of students.
[0600] The "analysis method" is a function that evaluates students' learning progress and emotional state based on acquired video information.
[0601] "Generation means" refers to a function that generates individual learning information based on the learning history and analysis results.
[0602] "Means of provision" refers to a function or device for providing students with generated learning information.
[0603] A "learning tool" is a function that accumulates feedback information and new learning information to optimize the entire system.
[0604] A "means of response" refers to a function or process for providing an immediate and appropriate answer to a student's question.
[0605] "Analysis means" refers to functions and processes for analyzing the user's emotional state and adjusting the teaching method accordingly.
[0606] "Means of providing feedback" refers to a function that notifies users of analysis results and emotional states in real time.
[0607] This system is designed to provide efficient learning support in educational settings. Its main components include terminals, servers, and users (primarily teachers).
[0608] The terminal is installed in the classroom and is a device for acquiring students' audio and video in real time. It incorporates an audio acquisition device and a camera to reliably capture students' audio and video information. This data is transmitted to a server via the internet or a local network.
[0609] The server plays a central role in data processing. Audio data is converted into text data using speech recognition software (e.g., a general-purpose speech recognition API). Next, natural language processing is performed using a generative AI model to analyze the content of the students' questions. Based on the analysis results, a prompt such as "Please create an explanation for the student's question" is input to the generative AI model to obtain an appropriate response.
[0610] Meanwhile, the video data is analyzed by emotion recognition software utilizing deep learning technology (e.g., open-source emotion recognition libraries). This allows for the evaluation of students' emotional states and concentration levels from their facial expressions and movements. Furthermore, the server generates personalized learning content based on the students' learning history and evaluation results, and provides it to the students via their devices.
[0611] The user, acting as a teacher, can analyze their own emotional state using an emotion engine. The device recognizes the teacher's voice and facial expressions, and sends the results to the server, which then adjusts the teaching method. Furthermore, the server has a function to provide feedback on the teacher's emotional state, supporting improved teaching quality and stress management.
[0612] For example, if a student asks a question such as, "Please explain photosynthesis in plants," the device captures the audio and analyzes it on a server. Based on the analysis, a generating AI model provides a specific and clear explanation, which is then immediately answered for the student. In addition, if a teacher's stress level is high, a warning alert is issued, and appropriate measures are taken depending on the cause.
[0613] In this way, the entire system supports individual student learning and effectively assists teachers in their instruction, thereby contributing to an improvement in the quality of education.
[0614] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0615] Step 1:
[0616] The device acquires audio and video of students in the classroom. The input consists of student speech and actions. Audio is collected via a microphone device, and video is captured by a camera. This data is then appropriately packaged for transmission to the server.
[0617] Step 2:
[0618] The server receives audio data transmitted from the terminal. The input here is the student's voice signal. The server uses speech recognition software to convert this audio signal into text data. The output is text representing what the student said.
[0619] Step 3:
[0620] The server analyzes the converted text data. A generative AI model is used for this, with the converted text as input. It utilizes natural language processing techniques to understand the meaning of the text and identifies the student's question as output. Specifically, it uses prompts to instruct the generative AI model to generate answers appropriate to the question.
[0621] Step 4:
[0622] The server generates answers to student questions. The input consists of the previously identified question and associated prompt. A generative AI model is used to generate appropriate answers, obtaining the answer text as output. This answer is then ready to be presented to the student.
[0623] Step 5:
[0624] The device provides feedback to the student based on the answers provided by the server. The input is the answer data sent from the server. The device conveys this data to the student as audio or text display.
[0625] Step 6:
[0626] The server analyzes video data of students sent from their terminals. In this step, deep learning is used to analyze students' facial expressions and movements. The input is video data, and the output generates evaluation results of the students' emotional state and concentration level.
[0627] Step 7:
[0628] The server records students' learning progress based on evaluation results and generates individualized learning content as needed. The input consists of emotional state and learning history data, and the output generated based on this data is personalized learning content. This enables more effective learning support for students.
[0629] Step 8:
[0630] The terminal acquires the emotional state of the user (teacher) and sends it to the server. The input here is the teacher's voice and facial expression information. The server uses an emotion engine to analyze this data and identifies the teacher's emotional state as output.
[0631] Step 9:
[0632] The server provides feedback tailored to the teacher's emotional state. The input is analyzed emotional data, and the output includes suggestions for adjusting teaching methods for the teacher. In this step, the teacher receives real-time feedback on their own emotional state and is prompted to take necessary actions.
[0633] (Application Example 2)
[0634] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0635] The problem that this invention aims to solve is to provide an advanced and flexible system for individually supporting students' learning in educational settings. In particular, there is a need to create a more effective educational environment by optimizing learning content while taking into account students' emotional states and levels of concentration, and by monitoring teachers' emotional states.
[0636] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0637] In this invention, the server includes a feedback means for collecting student responses and providing emotional feedback to teachers, an analysis means for analyzing the emotional state of educators and supporting improvement of instruction, and a means for analyzing emotions from video data and providing optimized educational support. This enables learning support tailored to each student, improvement of teachers' instructional skills, and effective stress management.
[0638] "Means of acquisition" refers to devices or methods for collecting students' audio or video information.
[0639] "Recognition means" refers to devices or methods that have the function of analyzing acquired audio information and converting it into text data.
[0640] "Processing means" refers to devices or methods for analyzing text data to identify the content of students' questions and derive appropriate responses.
[0641] "Analysis means" refers to devices or methods for analyzing video information and evaluating students' learning progress and emotional state.
[0642] "Generation means" refers to a device or method for creating individualized learning content based on students' learning history and analysis results.
[0643] "Means of delivery" refers to the devices and methods used to deliver the generated learning content to students.
[0644] A "feedback tool" is a device or method that collects student responses and provides emotional feedback to the teacher.
[0645] A "learning tool" is a device or method for accumulating feedback data and new learning data to optimize the entire system.
[0646] "Detection means" refers to devices or methods for identifying students' poor health or abnormal behavior and generating alerts in advance.
[0647] A "response mechanism" refers to a device or method for providing immediate answers to students' questions.
[0648] This invention provides student learning support through a system using an AI robot teacher for educational purposes. The system mainly consists of a server, terminals, and various analysis means.
[0649] The server receives audio and video information from the student's device and processes this information. The audio information is converted into text data using the Google Cloud Speech-to-Text API, and OpenAI's GPT model is used for natural language processing. At this stage, the server analyzes the student's question and generates an appropriate answer.
[0650] The server uses the Microsoft Azure Emotion API to perform emotion analysis on the video data. This allows the server to evaluate students' learning progress and emotional state, and then provide personalized learning content and support.
[0651] Furthermore, the feedback system notifies teachers of their emotional state in real time, supporting the improvement of their teaching skills. The feedback function includes an interface that visually shows how the teacher's own emotions are affecting their lessons.
[0652] For example, if a student asks, "I don't know how to find the area of this triangle," the system recognizes the voice and provides a real-time answer such as, "The area of a triangle can be calculated by base × height ÷ 2. For example, if the base is 5 cm and the height is 4 cm, the area is 10 square centimeters." Furthermore, support is provided based on the emotional state of the student and teacher.
[0653] Examples of prompts to input into a generative AI model:
[0654] A student asked, "How do you find the area of a triangle?" Please provide an educational and easy-to-understand explanation to answer this question.
[0655] This allows the entire system to provide more effective educational support based on two-way feedback between students and teachers.
[0656] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0657] Step 1:
[0658] The device acquires student audio and video information. The input is the student's raw audio and video. The device uses its built-in camera and microphone to collect this data in real time and send it to the server. The output is the transmission of data to the server.
[0659] Step 2:
[0660] The server converts the received audio information into text data using the Google Cloud Speech-to-Text API. The input is the student's audio data. The server uses speech recognition technology to analyze the audio data and generate text data. The output is the student's question in text form.
[0661] Step 3:
[0662] The server inputs text data into an OpenAI GPT model, analyzes the question, and generates an answer. The input consists of student questions transcribed from speech. Natural language processing techniques are used to process the data and generate appropriate answers. The output is the answer text that can be provided to the student.
[0663] Step 4:
[0664] The server simultaneously analyzes the received video data using the Microsoft Azure Emotion API to evaluate the students' emotional state. The input is the students' video data. Deep learning technology is used to recognize emotions and determine the students' emotional state and level of concentration. The output is data indicating the students' emotional state.
[0665] Step 5:
[0666] The server creates a plan to provide individualized learning content and support based on the generated responses and emotional state data. The input is the generated responses and emotional state data. The server then develops an individualized learning plan by comparing it with the student's learning history and other factors. The output is the learning content provided to the student.
[0667] Step 6:
[0668] The terminal delivers the provided learning content to the students. The input is learning content data sent from the server. The terminal displays the content to the students via its display and provides learning support. The output is learning information displayed to the students.
[0669] Step 7:
[0670] The server provides visual feedback to the user (teacher) regarding the student's learning progress and the teacher's own emotional state using feedback mechanisms. The input consists of student learning progress and teacher's emotional data. Based on this data, the server generates feedback information for the teacher. The output is the feedback information provided to the teacher.
[0671] Through these steps, the server, terminal, and user work together to provide detailed, personalized learning support for each student.
[0672] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0673] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0674] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0675] [Fourth Embodiment]
[0676] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0677] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0678] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0679] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0680] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0681] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0682] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0683] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0684] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0685] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0686] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0687] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0688] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0689] This invention is a learning support system using an AI robot teacher in educational settings, aiming to reduce the burden on teachers and improve the quality of individualized instruction for students. This system mainly consists of terminals and a server.
[0690] The terminals are installed in the classrooms and are equipped with microphones to capture student voices in real time, and cameras to monitor students' facial expressions and behavior. As a result, audio and video information is acquired simultaneously and transmitted to the server.
[0691] The server uses a speech recognition engine to convert the audio information sent from the terminal into text data. At this stage, questions and requests are extracted from the audio. Furthermore, natural language processing is used to analyze the text data and understand the intent behind the student's questions. This analysis generates appropriate answers and learning content.
[0692] The server uses deep learning algorithms to analyze the video data, analyzing students' emotional states and concentration levels. This allows teachers to identify students who require individual attention and to evaluate the overall learning environment of the class. Furthermore, based on this data, an anomaly detection system immediately generates an alert if it detects any student distress or abnormal behavior, notifying the teacher.
[0693] Furthermore, the server uses a generative AI model to generate personalized learning content based on students' past learning history and progress data. This content is adjusted in difficulty and content according to the student's level of understanding, providing optimal learning materials tailored to individual needs.
[0694] Teachers, as users, can utilize the data and reports provided by their devices to understand their students' situations and develop effective teaching plans. This frees teachers from the detailed responses required in daily lessons, allowing them to dedicate more time and effort to individualized instruction.
[0695] Thus, the AI robot teacher system of the present invention can contribute to improving the efficiency and quality of educational settings by utilizing speech recognition and video analysis to automate student learning support.
[0696] The following describes the processing flow.
[0697] Step 1:
[0698] The device uses a microphone to capture audio information emitted by students in the classroom, while simultaneously capturing students' facial expressions and movements with a camera. This allows for the collection of audio and video data in real time.
[0699] Step 2:
[0700] The device sends the collected audio data to the server. The server receives this data and converts it into text data using a speech recognition engine. The recognized text is then processed to facilitate smooth responses to the teacher's questions.
[0701] Step 3:
[0702] The server analyzes the converted text data using natural language processing techniques to understand the intent and content of the questions. Based on this analysis, it generates answers to help students understand the questions.
[0703] Step 4:
[0704] The terminal sends video data to the server for analysis. The server uses deep learning technology to analyze the students' facial expressions and movements to evaluate their emotional state and level of concentration.
[0705] Step 5:
[0706] The server uses anomaly detection methods based on the analysis results to detect student malaise or abnormal behavior. If detected, an alert is immediately generated and notified to the user.
[0707] Step 6:
[0708] The server references the student's past learning history and generates optimized, personalized learning content using a generative AI model. The generated content is customized to the student's current level of understanding and progress.
[0709] Step 7:
[0710] The device provides students with generated learning content and receives real-time feedback from students through the device. This feedback is used to generate the next learning content.
[0711] Step 8:
[0712] Users can use real-time data and reports provided by their devices to understand students' learning progress and plan effective instruction. Teachers can more easily provide instruction based on quantitative data.
[0713] (Example 1)
[0714] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0715] There is a need to reduce the burden on teachers in educational settings and improve the quality of individualized instruction for students. Traditional methods have made it difficult to effectively respond to students' learning situations and individual needs, and it has also been difficult to grasp the condition of individual students in large groups. Furthermore, there are insufficient means to detect and respond to student distress or abnormal behavior at an early stage, so these challenges need to be addressed.
[0716] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0717] In this invention, the server includes a recording device for acquiring audio, a conversion device for converting audio information into text, and an analysis device for analyzing text information and extracting intent. This enables effective individualized instruction tailored to the student's learning situation and needs.
[0718] A "recording device" is a device used to collect audio information from students.
[0719] A "conversion device" is a device used to convert acquired audio information into text data.
[0720] An "analysis device" is a device that examines text data in detail to extract the student's intentions and the content of their questions.
[0721] A "filming device" is a device used to acquire video information of students.
[0722] A "judgment device" is a device that analyzes video information of students to evaluate their learning progress and emotional state.
[0723] A "generation device" is a device that creates individualized learning materials based on students' learning history and analysis results.
[0724] A "distribution device" is a device used to supply students with the generated teaching materials.
[0725] A "memory device" is a device used to store learning data and optimize the system.
[0726] A "detection device" is a device that detects students' poor health or abnormal behavior and generates alerts in advance.
[0727] A "response device" is a device that provides immediate answers to students' questions.
[0728] This invention is an AI system intended to support classroom instruction in educational settings. It mainly consists of a terminal installed in the classroom and a server for data processing. The terminal includes a high-sensitivity microphone for acquiring student voice information and a high-resolution camera for recording student video information.
[0729] The server uses a speech recognition engine to convert the audio information sent from the terminal into text data. A common example of software used in this speech recognition process is a cloud-based speech recognition service. The text data is then analyzed using natural language processing techniques to clarify the intent behind the student's questions and requests. For example, a natural language processing library might be used for this process.
[0730] Regarding video data, the server utilizes deep learning to evaluate students' emotional states and concentration levels. The algorithms used in this evaluation process include open-source machine learning frameworks. Furthermore, if any distress or abnormal behavior is detected based on the analysis results, an alert is immediately generated and notified to the teacher's terminal.
[0731] Furthermore, the server analyzes students' past learning history and uses a generative AI model to generate personalized learning materials. These materials are adjusted according to the student's level of understanding, providing a learning experience tailored to individual learning needs.
[0732] For example, if a student says, "I don't understand the meaning of differentiation," the system can analyze their intent, generate and provide learning materials that include the basic concepts of differentiation. An example of a prompt from this system is: "When a student asks a question about something they don't understand during class, analyze their intent and generate appropriate learning content."
[0733] As described above, the present invention reduces the burden on teachers and makes it possible to efficiently provide individualized instruction tailored to each student.
[0734] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0735] Step 1:
[0736] The device uses a microphone to capture audio from students in the classroom. The input is the student's voice, which is received by the device. Specifically, during class, asking a question like, "How do you solve this equation?" collects the audio data. The output is the collected audio data, which is sent to the server.
[0737] Step 2:
[0738] The server processes the transmitted audio data using a speech recognition engine and converts it into text data. The audio data received from the terminal is used as input. Specifically, the speech recognition engine recognizes the audio "How do you solve this equation?" and generates the corresponding text "How do you solve this equation?". The generated text data is obtained as output.
[0739] Step 3:
[0740] The server analyzes the generated text data using natural language processing techniques to extract the intent behind the student's question. Text data obtained from speech recognition is used as input. Specifically, the natural language processing technique identifies the intent, "I want to know how to solve this equation." The output reveals the intent behind the question.
[0741] Step 4:
[0742] The server processes video data transmitted from the terminal using a deep learning algorithm for emotion analysis and concentration assessment. The input is the video data received from the terminal. Specifically, it analyzes the video and detects an emotional state such as "confused" from the student's facial expressions and posture. The output is the analyzed emotional state information.
[0743] Step 5:
[0744] The server uses a generative AI model to generate learning content tailored to each student, based on the intent and emotional state of the question. The input includes the student's past learning history and analyzed data. Specifically, the AI model automatically creates materials including "basic equations and how to solve them." The output is personalized learning content.
[0745] Step 6:
[0746] The user, the teacher, reviews the generated learning content and analysis results, and provides feedback to students based on them. The inputs include content provided by the server and student status information. Specifically, the teacher utilizes the information from the system to provide individualized instruction to students during class. The output is the provision of appropriate guidance and support to students.
[0747] (Application Example 1)
[0748] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0749] To improve factory production efficiency, it is necessary to accurately understand workers' actions and emotional states and provide them with the most appropriate work procedures in a timely manner. However, conventional systems have difficulty analyzing workers' actions and emotions in real time, and are unable to detect potential hazards in advance and respond appropriately, resulting in insufficient improvements in work efficiency and safety.
[0750] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0751] In this invention, the server includes voice acquisition means for acquiring worker movement information, video analysis means for analyzing worker video information and evaluating production efficiency, and anomaly detection means for alerting inappropriate movements in advance. This makes it possible to grasp the worker's movements and emotions in real time, provide optimal work procedures, and avoid potential hazards.
[0752] "Workers" refer to people who assemble parts or process products within a factory, and their actions and emotional states are the targets for management and optimization.
[0753] "Motion information" refers to data about the movements and work attitude of workers, and is real-time information acquired through sensors such as cameras.
[0754] A "voice acquisition means" is a device that has the function of collecting the voice emitted by a worker as digital data using a microphone or other device.
[0755] "Speech recognition means" refers to a technology that has the function of converting acquired speech data into text data and analyzes the content of what the worker says as text information.
[0756] "Natural language processing means" refers to technologies that analyze text data generated by speech recognition means and understand its context and intent.
[0757] "Video acquisition means" refers to a device that has the function of recording the actions of workers as video using a camera or similar device, and is used for analyzing their movements and emotions.
[0758] "Video analysis means" refers to technology that processes acquired video data to evaluate the productivity and emotional state of workers.
[0759] A "content generation method" is a system that has the function of generating individually optimized work content based on the results of analysis of the worker's actions and emotions.
[0760] A "content delivery means" is a device that has the function of communicating the generated optimization work content to workers and supporting them in carrying out their tasks.
[0761] "Learning methods" refer to technologies that accumulate feedback data and newly acquired work data to improve the accuracy and efficiency of the system.
[0762] An "anomaly detection method" is a technology that detects inappropriate actions or potential hazards by workers in real time and issues an immediate alert.
[0763] A "real-time response system" is a technology that provides optimal work procedures and advice in response to the worker's instructions and situation.
[0764] To implement this invention, a system is configured in which a server, terminal, and user cooperate to evaluate and optimize the worker's actions in real time. The server continuously acquires data from the work site by utilizing voice acquisition means and video acquisition means to collect worker action information. The voice acquisition means uses a microphone to capture the worker's speech and converts it into text data using speech recognition means. Natural language processing software analyzes the worker's intentions and instructions from the text data.
[0765] Simultaneously, the video data collected by the video acquisition system is processed using video analysis tools that apply deep learning algorithms to evaluate the workers' actions and emotions. This makes it possible to grasp production efficiency and the workers' psychological state in real time.
[0766] Based on the analysis results, the server uses a generation AI model to create individually optimized work instructions in the content generation means, and provides appropriate work instructions to the user (worker) via the content provision means. In addition, if the anomaly detection means detects inappropriate operation or potential danger, an alert is immediately generated and notified to the worker and supervisor.
[0767] The learning method analyzes accumulated feedback data and newly acquired work data to continuously optimize the system. The hardware used in this process includes professional cameras, high-sensitivity microphones, and a high-performance computing platform located on a server as computing resources. Software such as OpenCV, Keras, and SpeechRecognition are employed.
[0768] As a concrete example, when a worker is assembling parts, the server detects the worker's anxious expression using video analysis, and a generative AI model generates encouraging comments and provides voice feedback through a content delivery system, such as "Please feel free to contact us if you have any problems." An example of a prompt message would be, "Please propose a new strategy to optimize the next automation process." In this way, the system plays a role in enhancing the smoothness and safety of work.
[0769] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0770] Step 1:
[0771] The server receives voice data from the worker, acquired via microphone, as input from the terminal. This voice data is then used as input for speech recognition, digital signal processing is performed, and it is converted into text data for output. Specifically, the characteristics of the voice waveform are analyzed, and recognition is performed using a language model.
[0772] Step 2:
[0773] The server uses the text data output in step 1 as input and processes it with natural language processing. Here, text analysis is performed to identify the intent behind the worker's instructions and questions, and data is output to determine actions based on that intent. Context and intent are interpreted using natural language processing technology.
[0774] Step 3:
[0775] The server acquires video information of workers as input via cameras connected to terminals. This data is then processed by a video analysis system, and a deep learning algorithm is used to evaluate production efficiency and the emotional state of the workers, outputting the evaluation results. Facial expressions and movement patterns are analyzed from the video data to infer psychological state and work speed.
[0776] Step 4:
[0777] The server uses the intent and evaluation results obtained in steps 2 and 3 as input to generate individually optimized work content using a generative AI model. This generated work content is output by a content generation means. At this stage, prompts are used to generate work procedures and advice as specific instructions.
[0778] Step 5:
[0779] The server provides the content generated in step 4 to the user (worker) via the content delivery means. The outputted work instructions and advice are transmitted to the worker as audio or video, supporting the worker in properly performing their tasks.
[0780] Step 6:
[0781] The server stores feedback data collected during the work process and newly acquired work data as input to its learning mechanism. This allows for long-term data analysis to optimize the system and improve its accuracy. Statistical analysis of the data and machine learning are used to improve the accuracy of future work instructions.
[0782] Step 7:
[0783] The server uses anomaly detection mechanisms to detect inappropriate worker behavior and potential hazards in real time. Based on this input, it immediately generates and outputs alerts, notifying the user (worker) and relevant supervisors. This ensures workplace safety and enables a rapid response.
[0784] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0785] This invention aims to achieve more effective learning support by combining an emotion engine with a system that utilizes AI robot teachers in educational settings. This system mainly consists of a terminal, a server, and emotion recognition means.
[0786] The terminal has the ability to acquire students' audio and video in real time within the classroom and transmits this data to a server. The server converts the received audio data into text data using speech recognition technology and analyzes the content of the questions using natural language processing technology. Based on these results, specific answers and individualized learning content are generated to support students' learning.
[0787] The server analyzes the video data using deep learning technology to evaluate students' emotional states and concentration levels based on their facial expressions and movements. This evaluation allows the system to track students' learning progress, and individualized support is provided as needed through content generation methods.
[0788] The emotion engine is used specifically to analyze the emotional state of the user, the teacher. This allows for feedback that helps maintain or improve the quality of instruction based on the teacher's own emotions.
[0789] Specifically, the emotion recognition system analyzes the teacher's facial expressions and tone of voice, and sends the emotion data to a server. The server has the function to slightly adjust the teacher's teaching methods based on this data. In addition, it supports self-emotional management by providing real-time visual feedback on the teacher's emotional state.
[0790] For example, if a student asks a question during class, the device captures the audio, which is quickly analyzed on the server to provide the student with an immediate and appropriate answer. Furthermore, emotion recognition technology is used to monitor the teacher's stress level, and if high stress is detected, an alert is sent to the teacher as a warning. This data is later used to optimize the entire system and also helps maintain the teacher's mental health.
[0791] Thus, this system aims to improve the accuracy of individualized instruction, reduce the burden on teachers, and enhance the quality of education by utilizing multifaceted data, including feedback from students.
[0792] The following describes the processing flow.
[0793] Step 1:
[0794] The device collects students' voices in the classroom using a microphone and simultaneously captures video data such as facial expressions and movements using a camera. This makes it possible to acquire audio and video information in real time.
[0795] Step 2:
[0796] The server receives the audio data sent from the device and converts it into text data through a speech recognition engine. This text conversion allows for a more detailed understanding of the student's questions.
[0797] Step 3:
[0798] The server analyzes the converted text data using natural language processing to identify the intent and purpose of the questions. Based on these analysis results, appropriate answers and additional guidance for students are formulated.
[0799] Step 4:
[0800] The server receives video data transmitted from the terminal and analyzes it using deep learning. Here, the emotional state and concentration level of students are evaluated from their facial expressions and gestures, allowing for real-time monitoring of their learning progress.
[0801] Step 5:
[0802] Based on the analysis results, the server uses content generation methods to create personalized learning content for each student. Optimized content is provided based on historical learning data and current understanding.
[0803] Step 6:
[0804] The system uses emotion recognition to obtain user (teacher) emotion data, which is then sent to a server for analysis to understand the user's emotional state. The server generates feedback to adjust teaching content and methods according to stress levels and emotional changes.
[0805] Step 7:
[0806] Based on feedback from the server, users can monitor their own emotions and the students' learning progress to create effective teaching plans. Real-time visualization of emotional data supports emotional management during instruction.
[0807] Step 8:
[0808] All feedback data is stored on the server and continuously used to optimize the entire system. This ensures continuous improvement in the quality and efficiency of education.
[0809] (Example 2)
[0810] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0811] In educational settings, individualized support tailored to each student's learning situation and emotional state is required, but in reality, this is difficult to achieve due to the heavy burden on teachers. Furthermore, teachers' well-being can negatively impact the quality of instruction. Additionally, prompt and accurate responses to student questions are necessary. A system that comprehensively addresses these challenges is needed.
[0812] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0813] In this invention, the server includes recognition means for converting student voice information into text data, processing means for analyzing the content of student questions, and analysis means for analyzing student video information and evaluating learning progress. This enables learning support tailored to each student's learning situation, and also improves the quality of instruction through analysis of the teacher's emotional state and feedback function. Furthermore, by providing a function to respond to student questions immediately, it supports the rapid progress of lessons.
[0814] "Acquisition means" refers to a device or function that acquires students' audio and video information in real time.
[0815] "Recognition means" refers to software or a device that converts acquired audio information into text data.
[0816] "Processing means" refers to the functions and processes for analyzing the content of students' questions from text data.
[0817] "Video acquisition means" refers to a device or function for acquiring video information of students.
[0818] The "analysis method" is a function that evaluates students' learning progress and emotional state based on acquired video information.
[0819] "Generation means" refers to a function that generates individual learning information based on the learning history and analysis results.
[0820] "Means of provision" refers to a function or device for providing students with generated learning information.
[0821] A "learning tool" is a function that accumulates feedback information and new learning information to optimize the entire system.
[0822] A "means of response" refers to a function or process for providing an immediate and appropriate answer to a student's question.
[0823] "Analysis means" refers to functions and processes for analyzing the user's emotional state and adjusting the teaching method accordingly.
[0824] "Means of providing feedback" refers to a function that notifies users of analysis results and emotional states in real time.
[0825] This system is designed to provide efficient learning support in educational settings. Its main components include terminals, servers, and users (primarily teachers).
[0826] The terminal is installed in the classroom and is a device for acquiring students' audio and video in real time. It incorporates an audio acquisition device and a camera to reliably capture students' audio and video information. This data is transmitted to a server via the internet or a local network.
[0827] The server plays a central role in data processing. Audio data is converted into text data using speech recognition software (e.g., a general-purpose speech recognition API). Next, natural language processing is performed using a generative AI model to analyze the content of the students' questions. Based on the analysis results, a prompt such as "Please create an explanation for the student's question" is input to the generative AI model to obtain an appropriate response.
[0828] Meanwhile, the video data is analyzed by emotion recognition software utilizing deep learning technology (e.g., open-source emotion recognition libraries). This allows for the evaluation of students' emotional states and concentration levels from their facial expressions and movements. Furthermore, the server generates personalized learning content based on the students' learning history and evaluation results, and provides it to the students via their devices.
[0829] The user, acting as a teacher, can analyze their own emotional state using an emotion engine. The device recognizes the teacher's voice and facial expressions, and sends the results to the server, which then adjusts the teaching method. Furthermore, the server has a function to provide feedback on the teacher's emotional state, supporting improved teaching quality and stress management.
[0830] For example, if a student asks a question such as, "Please explain photosynthesis in plants," the device captures the audio and analyzes it on a server. Based on the analysis, a generating AI model provides a specific and clear explanation, which is then immediately answered for the student. In addition, if a teacher's stress level is high, a warning alert is issued, and appropriate measures are taken depending on the cause.
[0831] In this way, the entire system supports individual student learning and effectively assists teachers in their instruction, thereby contributing to an improvement in the quality of education.
[0832] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0833] Step 1:
[0834] The device acquires audio and video of students in the classroom. The input consists of student speech and actions. Audio is collected via a microphone device, and video is captured by a camera. This data is then appropriately packaged for transmission to the server.
[0835] Step 2:
[0836] The server receives audio data transmitted from the terminal. The input here is the student's voice signal. The server uses speech recognition software to convert this audio signal into text data. The output is text representing what the student said.
[0837] Step 3:
[0838] The server analyzes the converted text data. A generative AI model is used for this, with the converted text as input. It utilizes natural language processing techniques to understand the meaning of the text and identifies the student's question as output. Specifically, it uses prompts to instruct the generative AI model to generate answers appropriate to the question.
[0839] Step 4:
[0840] The server generates answers to student questions. The input consists of the previously identified question and associated prompt. A generative AI model is used to generate appropriate answers, obtaining the answer text as output. This answer is then ready to be presented to the student.
[0841] Step 5:
[0842] The device provides feedback to the student based on the answers provided by the server. The input is the answer data sent from the server. The device conveys this data to the student as audio or text display.
[0843] Step 6:
[0844] The server analyzes video data of students sent from their terminals. In this step, deep learning is used to analyze students' facial expressions and movements. The input is video data, and the output generates evaluation results of the students' emotional state and concentration level.
[0845] Step 7:
[0846] The server records students' learning progress based on evaluation results and generates individualized learning content as needed. The input consists of emotional state and learning history data, and the output generated based on this data is personalized learning content. This enables more effective learning support for students.
[0847] Step 8:
[0848] The terminal acquires the emotional state of the user (teacher) and sends it to the server. The input here is the teacher's voice and facial expression information. The server uses an emotion engine to analyze this data and identifies the teacher's emotional state as output.
[0849] Step 9:
[0850] The server provides feedback tailored to the teacher's emotional state. The input is analyzed emotional data, and the output includes suggestions for adjusting teaching methods for the teacher. In this step, the teacher receives real-time feedback on their own emotional state and is prompted to take necessary actions.
[0851] (Application Example 2)
[0852] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0853] The problem that this invention aims to solve is to provide an advanced and flexible system for individually supporting students' learning in educational settings. In particular, there is a need to create a more effective educational environment by optimizing learning content while taking into account students' emotional states and levels of concentration, and by monitoring teachers' emotional states.
[0854] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0855] In this invention, the server includes a feedback means for collecting student responses and providing emotional feedback to teachers, an analysis means for analyzing the emotional state of educators and supporting improvement of instruction, and a means for analyzing emotions from video data and providing optimized educational support. This enables learning support tailored to each student, improvement of teachers' instructional skills, and effective stress management.
[0856] "Means of acquisition" refers to devices or methods for collecting students' audio or video information.
[0857] "Recognition means" refers to devices or methods that have the function of analyzing acquired audio information and converting it into text data.
[0858] "Processing means" refers to devices or methods for analyzing text data to identify the content of students' questions and derive appropriate responses.
[0859] "Analysis means" refers to devices or methods for analyzing video information and evaluating students' learning progress and emotional state.
[0860] "Generation means" refers to a device or method for creating individualized learning content based on students' learning history and analysis results.
[0861] "Means of delivery" refers to the devices and methods used to deliver the generated learning content to students.
[0862] A "feedback tool" is a device or method that collects student responses and provides emotional feedback to the teacher.
[0863] A "learning tool" is a device or method for accumulating feedback data and new learning data to optimize the entire system.
[0864] "Detection means" refers to devices or methods for identifying students' poor health or abnormal behavior and generating alerts in advance.
[0865] A "response mechanism" refers to a device or method for providing immediate answers to students' questions.
[0866] This invention provides student learning support through a system using an AI robot teacher for educational purposes. The system mainly consists of a server, terminals, and various analysis means.
[0867] The server receives audio and video information from the student's device and processes this information. The audio information is converted into text data using the Google Cloud Speech-to-Text API, and OpenAI's GPT model is used for natural language processing. At this stage, the server analyzes the student's question and generates an appropriate answer.
[0868] The server uses the Microsoft Azure Emotion API to perform emotion analysis on the video data. This allows the server to evaluate students' learning progress and emotional state, and then provide personalized learning content and support.
[0869] Furthermore, the feedback system notifies teachers of their emotional state in real time, supporting the improvement of their teaching skills. The feedback function includes an interface that visually shows how the teacher's own emotions are affecting their lessons.
[0870] For example, if a student asks, "I don't know how to find the area of this triangle," the system recognizes the voice and provides a real-time answer such as, "The area of a triangle can be calculated by base × height ÷ 2. For example, if the base is 5 cm and the height is 4 cm, the area is 10 square centimeters." Furthermore, support is provided based on the emotional state of the student and teacher.
[0871] Examples of prompts to input into a generative AI model:
[0872] A student asked, "How do you find the area of a triangle?" Please provide an educational and easy-to-understand explanation to answer this question.
[0873] This allows the entire system to provide more effective educational support based on two-way feedback between students and teachers.
[0874] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0875] Step 1:
[0876] The device acquires student audio and video information. The input is the student's raw audio and video. The device uses its built-in camera and microphone to collect this data in real time and send it to the server. The output is the transmission of data to the server.
[0877] Step 2:
[0878] The server converts the received audio information into text data using the Google Cloud Speech-to-Text API. The input is the student's audio data. The server uses speech recognition technology to analyze the audio data and generate text data. The output is the student's question in text form.
[0879] Step 3:
[0880] The server inputs text data into an OpenAI GPT model, analyzes the question, and generates an answer. The input consists of student questions transcribed from speech. Natural language processing techniques are used to process the data and generate appropriate answers. The output is the answer text that can be provided to the student.
[0881] Step 4:
[0882] The server simultaneously analyzes the received video data using the Microsoft Azure Emotion API to evaluate the students' emotional state. The input is the students' video data. Deep learning technology is used to recognize emotions and determine the students' emotional state and level of concentration. The output is data indicating the students' emotional state.
[0883] Step 5:
[0884] The server creates a plan to provide individualized learning content and support based on the generated responses and emotional state data. The input is the generated responses and emotional state data. The server then develops an individualized learning plan by comparing it with the student's learning history and other factors. The output is the learning content provided to the student.
[0885] Step 6:
[0886] The terminal delivers the provided learning content to the students. The input is learning content data sent from the server. The terminal displays the content to the students via its display and provides learning support. The output is learning information displayed to the students.
[0887] Step 7:
[0888] The server provides visual feedback to the user (teacher) regarding the student's learning progress and the teacher's own emotional state using feedback mechanisms. The input consists of student learning progress and teacher's emotional data. Based on this data, the server generates feedback information for the teacher. The output is the feedback information provided to the teacher.
[0889] Through these steps, the server, terminal, and user work together to provide detailed, personalized learning support for each student.
[0890] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0891] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0892] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0893] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0894] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0895] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0896] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0897] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0898] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0899] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0900] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0901] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0902] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0903] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0904] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0905] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0906] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0907] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0908] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0909] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0910] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.
[0911] The following is further disclosed regarding the embodiments described above.
[0912] (Claim 1)
[0913] A means of acquiring audio information of students,
[0914] A speech recognition means that converts the student's voice information into text data,
[0915] A natural language processing means that analyzes the aforementioned text data to identify the content of the student's question,
[0916] A means of acquiring video information of students,
[0917] A video analysis means for analyzing the video information of the aforementioned students and evaluating their learning status and emotional state,
[0918] A content generation means that generates individual learning content based on students' learning history and analysis results,
[0919] A content provision means for providing the generated learning content to students,
[0920] A system that includes a learning method for accumulating feedback data and new learning data to optimize the system.
[0921] (Claim 2)
[0922] The system according to claim 1, comprising an anomaly detection means for detecting student malaise or abnormal behavior and generating an alert in advance.
[0923] (Claim 3)
[0924] The system according to claim 1, comprising a real-time response means for providing immediate answers to students' questions.
[0925] "Example 1"
[0926] (Claim 1)
[0927] A recording device for acquiring sound,
[0928] A conversion device for converting audio information into text,
[0929] An analysis device for analyzing the aforementioned textual information and extracting intent,
[0930] A camera for acquiring images,
[0931] A judgment device for analyzing the aforementioned video information and evaluating the situation and emotions,
[0932] A generation device for generating individual learning materials based on learning history and analysis results,
[0933] A providing device for providing the generated teaching materials,
[0934] A system that includes a memory device for accumulating learning data and improving knowledge.
[0935] (Claim 2)
[0936] The system according to claim 1, further comprising a detection device for detecting malfunctions or abnormalities and issuing warnings in advance.
[0937] (Claim 3)
[0938] The system according to claim 1, comprising a response device for providing immediate answers to questions.
[0939] "Application Example 1"
[0940] (Claim 1)
[0941] A voice acquisition means for acquiring information on the worker's actions,
[0942] A speech recognition means that converts the operator's action information into text data,
[0943] A natural language processing means that analyzes the aforementioned text data to identify the content of the worker's instructions,
[0944] A video acquisition means for acquiring video information of workers,
[0945] A video analysis means for analyzing the video information of the aforementioned worker and evaluating production efficiency and emotional state,
[0946] A content generation means that generates individual work content based on the worker's work history and analysis results,
[0947] A content provision means for providing the generated work content to the worker,
[0948] A system that includes a learning method to optimize the system by accumulating feedback data and new work data.
[0949] (Claim 2)
[0950] The system according to claim 1, comprising an anomaly detection means for detecting inappropriate actions or potential hazards by workers and generating alerts in advance.
[0951] (Claim 3)
[0952] The system according to claim 1, comprising a real-time response means that immediately provides the optimal procedure in response to instructions from an operator.
[0953] "Example 2 of combining an emotion engine"
[0954] (Claim 1)
[0955] A means of acquiring student voice information,
[0956] A recognition means for converting the aforementioned student's voice information into text data,
[0957] A processing means for analyzing the aforementioned text data to identify the content of the student's question,
[0958] A means of acquiring video information of students,
[0959] An analysis means for analyzing the video information of the aforementioned students and evaluating their learning status and emotional state,
[0960] A generation means for generating individual learning information based on students' learning history and analysis results,
[0961] A means for providing the generated learning information to students,
[0962] A learning method that accumulates feedback information and new learning information to optimize the system,
[0963] An analytical means that analyzes the user's emotional state and adjusts the instruction method,
[0964] A means of providing feedback to the user about their emotional state,
[0965] A system that includes a response mechanism to provide immediate answers to students' questions.
[0966] (Claim 2)
[0967] The system according to claim 1, comprising detection means for detecting student malaise or abnormal behavior and generating an alert in advance.
[0968] (Claim 3)
[0969] The system according to claim 1, which uses a generative AI model when providing the generated learning information.
[0970] "Application example 2 when combining with an emotional engine"
[0971] (Claim 1)
[0972] A means of acquiring student voice information,
[0973] A recognition means for converting the aforementioned student's voice information into text data,
[0974] A processing means for analyzing the aforementioned text data to identify the content of the student's question,
[0975] Methods for acquiring video information of students,
[0976] An analysis means for analyzing the video information of the aforementioned students and evaluating their learning status and emotional state,
[0977] A generation means for generating individual learning materials based on students' learning history and analysis results,
[0978] A means for providing the generated learning content to students,
[0979] A feedback system that collects student responses and provides emotional feedback to teachers,
[0980] An analytical tool to analyze the emotional state of educators and support improvement of teaching methods,
[0981] A method for analyzing emotions from video data and providing optimized educational support,
[0982] A system that includes a learning method for accumulating feedback data and new learning data to optimize the system.
[0983] (Claim 2)
[0984] The system according to claim 1, comprising detection means for detecting a student's poor condition or abnormal behavior and generating an alert in advance.
[0985] (Claim 3)
[0986] The system according to claim 1, comprising a response means for providing immediate answers to students' questions. [Explanation of symbols]
[0987] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of acquiring audio information of students, A speech recognition means that converts the student's voice information into text data, A natural language processing means that analyzes the aforementioned text data to identify the content of the student's question, A means of acquiring video information of students, A video analysis means for analyzing the video information of the aforementioned students and evaluating their learning status and emotional state, A content generation means that generates individual learning content based on students' learning history and analysis results, A content provision means for providing the generated learning content to students, A system that includes a learning method for accumulating feedback data and new learning data to optimize the system.
2. The system according to claim 1, comprising an anomaly detection means for detecting student malaise or abnormal behavior and generating an alert in advance.
3. The system according to claim 1, comprising a real-time response means for providing immediate answers to students' questions.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A