system

The learning support system addresses high costs and time constraints by capturing and analyzing learners' activities in real-time, providing personalized feedback through AI, and adjusting learning plans based on emotional states, enhancing learning efficiency.

JP2026036117APending Publication Date: 2026-03-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-20
Publication Date
2026-03-05

AI Technical Summary

Technical Problem

Conventional tutoring services face challenges such as high usage costs and time constraints, with tutors often unable to provide immediate support, and lack the capability to analyze learners' facial expressions and gaze to optimize learning support.

Method used

A learning support system that captures learners' activities using cameras, transmits video data to a server for analysis, generates personalized feedback using AI, and provides it in real-time through a tablet or smartphone, incorporating emotion recognition to adjust learning plans.

Benefits of technology

Enables high-quality, real-time, individually optimized learning support by analyzing facial expressions, movements, and emotional states, improving learning efficiency and adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026036117000001_ABST
    Figure 2026036117000001_ABST
Patent Text Reader

Abstract

Provide a system. A photographing means for photographing a learner's learning activity; a transmitting means for transmitting the video data captured by the imaging means to a server; an analysis means in the server for analyzing the received video data; a generation means for generating feedback appropriate for the learner based on the analysis result; The system includes a means for providing the generated feedback to the learner.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] While conventional tutoring services offer the advantage of providing individually optimized learning, they also face challenges such as high usage costs and time constraints. In particular, when learners need tutoring support, the tutors are often unable to respond immediately. The present invention aims to solve these challenges and realize a system that provides high-quality learning support anytime, anywhere. [Means for solving the problem]

[0005] The present invention provides a learning support system having the following configuration: First, a camera means is provided for capturing images of a learner's learning activities. Next, a transmission means is provided for transmitting the video data captured by the camera means to a server. The server side is provided with a generation means that uses an analysis means for analyzing the received video data and generates feedback appropriate for the learner based on the analysis results. Finally, a provision means is provided for providing the generated feedback to the learner. This system allows learners to receive individually optimized learning support in real time. Furthermore, the camera means has the function of detecting the learner's facial expressions, movements, and line of sight and estimating their concentration and fatigue levels, making it possible to provide more accurate feedback. Furthermore, by providing a learning progress management means that records learning progress and adjusts the learning plan as needed, long-term learning support is realized.

[0006] "Photography means" refers to a camera or other photographic device used to photograph learners' learning activities.

[0007] The "transmission means" is a communication module or program for transmitting the video data captured by the image capture means to the server.

[0008] "Analysis means" refers to software or algorithms that analyze the video data received by the server and determine the learner's learning status based on their facial expressions, movements, line of sight, etc.

[0009] The "generation means" refers to an artificial intelligence model or generation algorithm for generating feedback appropriate for the learner based on the results obtained by the analysis means.

[0010] The "means of provision" refers to a display device such as a tablet or smartphone or an application that provides the generated feedback to the learner.

[0011] "Level of concentration" is an index that indicates how focused a learner is on learning activities.

[0012] "Fatigue level" is an index that indicates the degree of fatigue a learner feels from learning activities.

[0013] A "learning progress management tool" is software or algorithms that record a learner's learning progress and adjust the learning plan based on the stored data. [Brief explanation of the drawings]

[0014] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0015] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0016] First, the terms used in the following description will be explained.

[0017] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0018] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0019] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0020] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0022] [First embodiment]

[0023] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0024] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0025] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0026] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0027] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0029] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0030] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0031] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0032] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0033] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0034] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0035] The present invention is a system for supporting learners' learning, specifically, a system that analyzes and supports learners' activities in real time using cameras and tablet devices installed in the learning environment. An embodiment of this system will be described in detail.

[0036] This system mainly comprises the following components: an imaging means, a transmitting means, an analyzing means, a generating means, and a providing means.

[0037] Filming method

[0038] Camera: A camera installed on each student's desk constantly records their learning activities. The camera is connected to a tablet device via Wi-Fi and captures video data in real time.

[0039] Transmission method

[0040] Device: The tablet device receives the video data captured by the camera, compresses it, and sends it to the server using a secure communication protocol. The video data is high resolution and can capture the learner's detailed movements and facial expressions.

[0041] Analysis means

[0042] Server: The server analyzes the received video data. Specific software algorithms are used to analyze the learner's facial expressions, movements, and changes in gaze. This analysis makes it possible to assess the learner's level of concentration and fatigue. Furthermore, by integrating and analyzing past learning data (class notes, past questions, mock test results, etc.), the server can identify the learner's strengths and weaknesses.

[0043] generation means

[0044] Server: Based on the analysis results, it generates optimized feedback for the learner. A generative AI model is used in this process to generate specific advice, supplementary materials, additional problems, etc. For example, if a learner is having trouble solving an equation, an explanation or practice problem will be provided to encourage review of that part.

[0045] Providing means

[0046] Device: The generated feedback is sent to a tablet device and displayed to the user. Specific study advice and the next task to tackle based on the analysis results are displayed on the tablet screen. The user can check this and continue studying.

[0047] Specific examples

[0048] For example, consider a user learning how to solve mathematical equations. A camera captures the user's learning, and the video data is sent to a server via a tablet device. The server analyzes the video data and detects that the user is making a calculation error in a particular step. It also identifies that the user has repeatedly made mistakes in this area based on past practice test results.

[0049] Based on this, the server generates detailed explanations for the user on "typical mistakes in solving equations and how to correct them" and sends them to the tablet device. The user can then review the explanations on the tablet and work on practice problems to correct the mistakes. This allows the user to receive specific, individually optimized learning support in real time.

[0050] As described above, the system of the present invention realizes high-quality learning support by capturing and analyzing learners' learning activities in real time and providing individually optimized feedback.

[0051] The processing flow will be explained below.

[0052] Step 1:

[0053] User: Starts the learning application on the tablet device and configures the camera connection settings.

[0054] Step 2:

[0055] Device: The tablet device connects to the tabletop camera via Wi-Fi and checks the camera's connection status.

[0056] Step 3:

[0057] Camera: The tabletop camera continuously captures the learner's learning activities, capturing images at a set frame rate and resolution.

[0058] Step 4:

[0059] Terminal: Compresses video data captured in real time and sends it to the server using a secure communication protocol.

[0060] Step 5:

[0061] Server: Receives compressed video data and temporarily stores it in a database.

[0062] Step 6:

[0063] Server: Analyzes the received video data using "Gemini (registered trademark)." Specifically, it analyzes the learner's facial expressions, movements, and changes in line of sight to evaluate the learner's level of concentration and fatigue.

[0064] Step 7:

[0065] Server: Based on the analysis results, the learner's strengths and weaknesses are identified, and past learning data (class notes, past questions, mock test results, etc.) is integrated as necessary for analysis.

[0066] Step 8:

[0067] Server: Based on the analysis results, a generative AI model generates optimized feedback for the learner, including specific advice, supplementary materials, and additional questions.

[0068] Step 9:

[0069] Server: Sends the generated feedback to the tablet device.

[0070] Step 10:

[0071] Device: The tablet device receives the feedback and displays it in the application's user interface.

[0072] Step 11:

[0073] User: Check the feedback displayed on the tablet and learn what to study next or what explanations to give.

[0074] Step 12:

[0075] Server: Records new progress data of learners and stores it in a database. Based on the progress, the learning plan and feedback are updated.

[0076] Step 13:

[0077] User: Check the learning progress and continue learning according to the plan. If necessary, repeat the cycle from step 1.

[0078] Example 1

[0079] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0080] Conventional learning support systems have difficulty analyzing learners' learning activities in detail in real time, making it difficult to provide individually optimized feedback in a timely manner. Furthermore, they lack the functionality to capture subtle changes in learners' facial expressions and gaze to evaluate their concentration and fatigue levels, making it difficult to optimally improve learners' performance.

[0081] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0082] In this invention, the server includes a photographing means, a transmitting means, an analyzing means, a generating means, and a providing means. This enables the learner's learning activities to be analyzed in real time and individually optimized feedback to be provided promptly. Specifically, the server uses a specific software algorithm to analyze the received video data and analyzes the learner's facial expressions, movements, and changes in gaze to evaluate the learner's level of concentration and fatigue. Based on the results, the server uses a generative AI model to generate feedback appropriate for the learner, and has a means for transmitting and displaying the feedback on the mobile information terminal. This allows the learner to receive optimal learning support in real time, improving the quality of their learning.

[0083] The "filming means" is a device that films the learning activities of the learner in real time and acquires video data.

[0084] The "transmission means" is a device or software that has the function of compressing the video data acquired by the image capture means and transmitting it to the server using a secure communication protocol.

[0085] "Analysis means" refers to a device or software that has the function of analyzing the video data received by the server using a specific software algorithm, and analyzing changes in the learner's facial expressions, movements, and gaze to evaluate their level of concentration and fatigue.

[0086] "Generation means" refers to a device or software that has the function of utilizing a generative AI model to automatically generate feedback optimized for the learner based on the analysis results obtained by the analysis means and past learning data.

[0087] The "providing means" is a device or software that has the function of transmitting the feedback generated by the generating means to the learner's mobile information terminal and displaying it.

[0088] A "generative AI model" is a model that uses artificial intelligence algorithms to generate appropriate feedback based on input data.

[0089] This system is constructed to analyze learners' learning activities in detail in real time and provide individually optimized feedback. The system consists of a capturing means, a transmitting means, an analyzing means, a generating means, and a providing means.

[0090] Filming method

[0091] Camera: A camera (e.g., a high-resolution webcam) installed on the student's desk continuously records the student's learning activities. The camera is connected to a tablet device via Wi-Fi and captures video data in real time.

[0092] Transmission method

[0093] Terminal: The tablet terminal receives the video data captured by the camera and automatically compresses it using an efficient video compression algorithm such as H.264. The compressed video data is then sent to the server using a secure communication protocol (e.g., HTTPS).

[0094] Analysis means

[0095] Server: The server analyzes the received video data. For example, image processing software algorithms such as OpenCV are used for the analysis. The analysis method detects changes in the learner's facial expressions, movements, and line of sight, and evaluates the learner's level of concentration and fatigue based on these.

[0096] generation means

[0097] Server: Based on the analysis results and past learning data, a generative AI model (e.g., a generative AI model) is used to generate optimized feedback for the learner. For example, if a learner is having trouble solving an equation, an explanation or exercise tailored to the content is generated.

[0098] Providing means

[0099] Device: The generated feedback is sent to the learner's tablet device and displayed to them. Specific study advice and the next task to be tackled based on the analysis results are displayed on the tablet screen.

[0100] Specific examples

[0101] For example, if a user is learning how to solve mathematical equations, a camera captures the user's learning activities and the video data is sent to a server via a tablet device. The server analyzes the video data and detects that the user is making a calculation error in a specific step. It also identifies that the user has repeatedly made mistakes in this area based on past practice test results. Based on this, the server generates detailed explanations for the user on "typical mistakes in solving equations and how to correct them" and sends them to the tablet device. The user then checks the explanations on the tablet and works on practice problems to correct their mistakes.

[0102] Prompt Sentence Examples

[0103] "Analyze where a user makes mistakes in solving math equations and generate specific explanations and practice questions for review."

[0104] "Calculate the learner's concentration and fatigue levels from video data and provide appropriate learning advice."

[0105] The above is an embodiment of the present invention. This system enables learners to receive optimal learning support in real time, thereby improving the quality of their learning.

[0106] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0107] Step 1:

[0108] Filming learning activities

[0109] Camera: A camera installed on each student's desk constantly records their learning activities, capturing in high resolution the student's facial expressions, movements, and gaze in detail.

[0110] Input: Real-time video of the learner

[0111] Output: High-resolution video data of learning activities

[0112] Step 2:

[0113] Video data compression and transmission

[0114] Terminal: The tablet terminal compresses the video data received from the camera using the H.264 algorithm, and then transmits the compressed data to the server using a secure communication protocol (e.g., HTTPS).

[0115] Input: High-resolution video of learning activities

[0116] Output: Compressed video data

[0117] Step 3:

[0118] Video data analysis

[0119] Server: Analyzes the received compressed video data. Image processing algorithms such as OpenCV are used for the analysis. The server analyzes the learner's facial expressions, movements, and changes in gaze to evaluate the learner's level of concentration and fatigue. It also compares this with past learning data.

[0120] Input: Compressed video data, past learning data

[0121] Output: Evaluation results of learner's concentration, fatigue, strengths and weaknesses

[0122] Step 4:

[0123] Generate feedback

[0124] Server: Based on the analysis results, the server sends a prompt to the generative AI model. The prompt includes the learner's situation and the feedback they need. The generative AI model generates specific feedback (explanations, practice questions, advice, etc.) based on the prompt.

[0125] Input: Analysis results, prompt text

[0126] Output: Generated feedback

[0127] Step 5:

[0128] Providing Feedback

[0129] Terminal: The feedback sent from the server is displayed on the tablet terminal. The user checks the feedback displayed on the tablet. Examples include "typical mistakes in solving equations and how to correct them."

[0130] Input: Generated feedback

[0131] Output: Feedback display on tablet device

[0132] Examples:

[0133] As an example, consider the case where a user is learning how to solve a mathematical equation.

[0134] Step 1: The camera captures the user's learning activity and captures it as high-resolution video data.

[0135] Step 2: The terminal compresses the video data using the H.264 algorithm and sends it to the server using the HTTPS protocol.

[0136] Step 3: The server analyzes the received video data using OpenCV to detect if the user is making an incorrect calculation at a specific step. It also integrates past learning data and evaluates the learner's strengths and weaknesses based on this.

[0137] Step 4: Based on the analysis results, the generative AI model is prompted to "analyze the user's mistakes in solving mathematical equations and generate specific explanations and practice questions for review." The generative AI model then generates specific feedback.

[0138] Step 5: The generated feedback is sent to the device, and "typical mistakes in solving equations and how to correct them" are displayed on the user's tablet. The user then tries learning again while looking at the results.

[0139] The above are the specific processing steps of the system of the present invention, which enable learners to receive optimal learning support in real time.

[0140] (Application example 1)

[0141] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0142] Conventional automated equipment operating in factories lacks real-time analysis functions to optimize efficiency and precision, and is unable to provide immediate feedback on improvements. As a result, waste and malfunctions tend to occur in the equipment's operation, leading to problems such as reduced productivity and poor quality.

[0143] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0144] In this invention, the server includes a photographing means for photographing the learning activities of the learner, a transmitting means for transmitting video data photographed by the photographing means to the server, an analyzing means for analyzing the received video data, a generating means for generating feedback appropriate for the learner based on the analysis results, a providing means for providing the generated feedback to the learner, an automation device operation analyzing means for analyzing the operation of automation devices operating in the factory and providing feedback on areas for improvement in efficiency and accuracy, and a display means for displaying the feedback. This makes it possible to monitor the operation of automation devices operating in the factory in real time and improve efficiency and accuracy based on the analysis results.

[0145] "Photographing means" refers to a camera or other image capturing device that captures the learning activities of learners and the operation of automated equipment in a factory.

[0146] The "transmission means" refers to a device or technology for transmitting the video data captured by the image capture means to the server, and is, for example, a device that uses a tablet terminal or a communication protocol.

[0147] The "analysis means" is a technology for analyzing the video data received in the server and evaluating the activities of the learners and the operation status of the automated devices.

[0148] "Generation means" refers to technology for generating appropriate feedback to learners and improvements to automated devices based on the analysis results.

[0149] "Providing means" refers to the devices and technologies used to display the generated feedback to the learner or operator.

[0150] "Automated equipment operation analysis means" is a technology for analyzing the operation of automated equipment operating in a factory and detecting areas for improvement in efficiency and accuracy.

[0151] "Display means" refers to a device or technology for visually presenting the analysis and generated feedback to the user, such as a tablet device or display.

[0152] "Real-time" refers to a time frame in which data is acquired, transmitted, analyzed, and feedback generated nearly simultaneously.

[0153] "Efficiency" is an index that indicates the degree to which automated equipment operating in a factory operates optimally and without waste.

[0154] "Accuracy" is a measure of the ability or degree to which an automated device accurately performs a required operation.

[0155] The present invention is a system for optimizing automated equipment in a factory by analyzing its operation and providing feedback on areas for improvement in efficiency and accuracy. This system is realized through the following steps.

[0156] System Program

[0157] This system mainly comprises an imaging means, a transmission means, an analysis means, a generation means, and a provision means.

[0158] Filming method

[0159] A camera is used as a recording medium. The camera records the operation of the automated equipment in real time and acquires the video data. The camera is connected to a tablet device via Wi-Fi or a wired connection.

[0160] Transmission method

[0161] The device receives the captured video data, compresses it, and transmits it to a server using a secure communication protocol. The high-resolution video data can capture even the smallest details of the automated equipment's operation.

[0162] Analysis means

[0163] The server analyzes the received video data using specific software algorithms (e.g., computer vision technology) to evaluate the efficiency and accuracy of the automated equipment's operation. This analysis identifies waste and malfunctions in the equipment's operation.

[0164] generation means

[0165] Based on the analysis results, feedback is generated to improve efficiency and accuracy. A generative AI model is used in this process to generate specific improvements and supplemental materials. For example, feedback such as "The movement of the right arm is slow and needs adjustment" is generated.

[0166] Providing means

[0167] The generated feedback is sent to the tablet device and displayed to the user. The feedback is provided in a visual format and includes specific instructions on how to improve.

[0168] Hardware and software used

[0169] Camera: Films the operation of automated equipment.

[0170] Tablet device: Receives video data captured by the camera and sends it to the server.

[0171] Server: Analyzes video data and generates feedback using computer vision techniques and generative AI models.

[0172] Specific examples

[0173] For example, if there is an automated machine that assembles a product, a camera will record its operation. The server will analyze the video data and identify the problem: "The right arm is moving slowly." Based on the analysis results, the generative AI model will generate feedback such as "Please adjust the right arm to move 10% faster," and send this to a tablet device. The user will then check this feedback on the tablet and adjust the machine's operation.

[0174] Prompt Sentence Examples

[0175] Configure your robot to analyze footage of its movements and identify areas for improvement in efficiency and accuracy. Analyze the movements in the footage and generate feedback that is user-friendly and specific.

[0176] This completes the detailed description of the embodiment of the present invention. This system makes it possible to monitor the operation of automated devices operating in a factory in real time and improve efficiency and accuracy based on the analysis results.

[0177] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0178] Step 1:

[0179] Camera image capture

[0180] Input: Factory automation equipment operation

[0181] Output: High-resolution video data

[0182] The camera captures the operation of the automated equipment in real time, capturing the operating status of the automated equipment and sending the video to a terminal as high-resolution video data.

[0183] Step 2:

[0184] Receive video data on the device

[0185] Input: High-resolution video data from a camera

[0186] Output: Compressed video data stored on the device

[0187] The device receives the video data captured by the camera and compresses it, making it ready to be sent to the server.

[0188] Step 3:

[0189] Compressed video data sent to server

[0190] Input: Compressed video data

[0191] Output: Video data sent to the server

[0192] The device transmits the compressed video data to the server using a secure communication protocol, with encryption techniques used to ensure data integrity and privacy.

[0193] Step 4:

[0194] Video data analysis by server

[0195] Input: Received video data

[0196] Output: Evaluation of the efficiency and accuracy of the operation as a result of the analysis

[0197] The server analyzes the received video data using computer vision techniques and performs calculations on the data to evaluate the operational status of the automated device (e.g., speed, accuracy of operation).

[0198] Step 5:

[0199] Feedback Generation

[0200] Input: Analysis results

[0201] Output: Feedback on areas for improvement in efficiency and accuracy

[0202] The server generates feedback based on the analysis results to improve efficiency and accuracy, using a generative AI model to generate specific improvements and supplementary materials.

[0203] Step 6:

[0204] Provide feedback to the device

[0205] Input: Generated feedback

[0206] Output: Feedback displayed on the terminal

[0207] The server sends the generated feedback to the device, which receives it and visually displays it to the user, allowing the user to confirm and implement specific improvements.

[0208] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0209] The present invention is a system for supporting learners in their learning, and is characterized by its ability to provide feedback based on the learner's emotional state by incorporating an emotion engine. An embodiment of this system will be described in detail.

[0210] This system mainly consists of the following components: a photographing means, a transmitting means, an analyzing means, a generating means, a providing means, and an emotion engine.

[0211] Filming method

[0212] Camera: A camera installed on each student's desk constantly records their learning activities. The camera is connected to a tablet device via Wi-Fi and captures video data in real time.

[0213] Transmission method

[0214] Terminal: The tablet terminal receives the video data captured by the camera, compresses it, and sends it to the server using a secure communication protocol.

[0215] Analysis means

[0216] Server: The server analyzes the received video data. Specific software algorithms are used to analyze the learner's facial expressions, movements, and changes in gaze. This analysis evaluates the learner's level of concentration and fatigue. It also integrates and analyzes past learning data (class notes, past questions, mock test results, etc.) to identify the learner's strengths and weaknesses.

[0217] Emotion Engine

[0218] Server: The emotion engine analyzes video and audio data to recognize the learner's emotional state. Specifically, it identifies the learner's emotions, such as joy, sadness, and stress, from facial expression and audio data. This emotional data is integrated with data obtained by the analysis means, enabling more accurate learning support.

[0219] generation means

[0220] Server: Generates optimized feedback for learners based on the analysis results and emotional data obtained by the emotion engine. A generative AI model is used in this process to generate specific advice, supplementary materials, additional questions, etc. For example, if a learner is feeling stressed, encouraging words and advice on how to relax are provided.

[0221] Providing means

[0222] Device: The generated feedback is sent to a tablet device and displayed to the user. Specific study advice and next tasks based on the analysis results and emotional state are displayed on the tablet screen. The user can confirm this and continue studying.

[0223] Specific examples

[0224] For example, consider a user learning how to solve a mathematical equation. A camera captures the user's learning, and the video data is sent to a server via a tablet device. The server analyzes the video data and detects that the user is making a calculation error at a specific step. Furthermore, the emotion engine recognizes from the user's facial expressions and voice that the user is feeling frustrated because they cannot solve the equation.

[0225] Based on this, the server generates detailed explanations for the user about "common mistakes in solving equations and how to correct them." It also adds encouraging words to the explanations, such as "This part is difficult, but try to stay calm. It's also recommended to take breaks." The generated feedback is sent to a tablet device, where the user can review it and work on practice problems to correct their mistakes. This allows the user to receive specific, emotionally appropriate feedback in real time.

[0226] As described above, the system of the present invention captures and analyzes learners' learning activities in real time and provides feedback that takes into account their emotional state, thereby realizing high-quality, individually optimized learning support.

[0227] The processing flow will be explained below.

[0228] Step 1:

[0229] User: Starts the learning application on the tablet device and configures the camera connection settings.

[0230] Step 2:

[0231] Device: The tablet device connects to the tabletop camera via Wi-Fi and checks the camera's connection status.

[0232] Step 3:

[0233] Camera: The tabletop camera continuously captures the learner's learning activities, capturing images at a set frame rate and resolution.

[0234] Step 4:

[0235] Terminal: Compresses video data captured in real time and sends it to the server using a secure communication protocol.

[0236] Step 5:

[0237] Server: Receives compressed video data and temporarily stores it in a database.

[0238] Step 6:

[0239] Server: Analyzes the received video data using the "analysis means." Specifically, it analyzes the learner's facial expressions, movements, and changes in line of sight to evaluate the learner's level of concentration and fatigue.

[0240] Step 7:

[0241] Server: Obtains past learning data (class notes, past questions, mock test results, etc.) from a database and integrates it with current analysis results to identify the learner's strengths and weaknesses.

[0242] Step 8:

[0243] Server: Using an "emotion engine," it analyzes video and audio data to recognize the learner's emotional state (happiness, sadness, stress, etc.). Specifically, it runs an algorithm that identifies emotions from facial expression and audio data.

[0244] Step 9:

[0245] Server: The "generator" generates optimal feedback based on the analysis results and emotional data. For example, if a calculation error is detected, it will include detailed instructions on how to correct the error and advice corresponding to the user's emotional state.

[0246] Step 10:

[0247] Server: Sends the generated feedback to the tablet device.

[0248] Step 11:

[0249] Device: The tablet device receives the feedback and displays it in the application's user interface.

[0250] Step 12:

[0251] User: Check the feedback displayed on the tablet and learn what to study next or what explanations to give. If necessary, follow the advice to relax.

[0252] Step 13:

[0253] Server: Records new progress and emotion data of the learner and stores it in a database. Based on the progress and emotion data, the learning plan and feedback content are updated.

[0254] Step 14:

[0255] User: Check their learning progress and emotional state, continue learning according to the plan, and repeat the cycle from step 1 if necessary.

[0256] Example 2

[0257] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0258] Conventional learning support systems can record learners' learning activities and provide feedback, but they have limitations in providing real-time support that takes into account the learner's emotional state. This makes it difficult to provide feedback appropriate to a learner's emotional state when they feel stressed or fatigued. Furthermore, there is a need for systems that go beyond simple data analysis to automatically identify a learner's strengths and weaknesses and provide individually optimized learning support based on that information.

[0259] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0260] In this invention, the server includes a photographing means for photographing the learning activities of the learner, a transmitting means for transmitting the video data photographed by the photographing means to the server, an analyzing means in the server for analyzing the received video data, a generating means for generating appropriate feedback based on the data and emotional data obtained by the analyzing means, a providing means for providing the generated feedback to the learner, an emotion recognition means for analyzing the emotional state of the learner from the video data and audio data, and a generating means using a generative AI model for integrating the obtained emotional data and data from the analyzing means to generate feedback optimized for the learner. This enables high-quality learning support that is individually optimized in real time and takes into account the emotional state of the learner.

[0261] "Photography means" refers to devices used to record learners' learning activities, such as cameras.

[0262] The "transmission means" refers to a device or function for transmitting the video data captured by the image capture means to the server.

[0263] "Analysis means" refers to software algorithms or programs used to analyze received video data and analyse changes in the learner's facial expressions, movements and gaze.

[0264] The "generation means" refers to a process or device that generates optimized feedback for the learner based on the data and emotional data obtained by the analysis means.

[0265] "Providing means" refers to a device or method for presenting the generated feedback to the learner.

[0266] "Emotion recognition means" refers to software or algorithms that analyze video and audio data to identify a learner's emotional state.

[0267] "Generative AI model" refers to an artificial intelligence model used to generate optimal feedback and advice for learners based on collected data.

[0268] "Feedback" refers to specific advice, supplementary materials, additional questions, words of encouragement, etc., generated by the analysis means or generation means and provided to the learner.

[0269] "Learning progress management means" refers to a process or device for recording a learner's progress and adjusting the learning plan based on the saved progress data.

[0270] The present invention is a system that comprehensively supports learners' learning activities, and its distinctive feature is that it incorporates the learner's emotional state. A specific embodiment of this system will be described. The main components include an image capture means, a transmission means, an analysis means, a generation means, a provision means, and an emotion recognition means. This makes it possible to provide learners with individually optimized feedback in real time.

[0271] First, when a user begins studying, a camera installed on the desk automatically starts operating. The camera captures the user's study activities in real time and continuously acquires video data. The camera is connected to a tablet device via Wi-Fi, and the captured video data is sent to the tablet device.

[0272] A camera connected to a terminal sends captured video data to the terminal. The terminal compresses the received video data and sends it to a server via a secure communication protocol (e.g., HTTPS). Data compression minimizes communication delays.

[0273] The server analyzes the received video data. Specifically, it uses OpenCV and other tools to analyze the user's facial expressions, movements, and changes in gaze. This allows it to evaluate the user's level of concentration and fatigue. It also integrates and analyzes past learning data (class notes, past questions, mock test results, etc.) to identify the user's strengths and weaknesses.

[0274] The emotion recognition means analyzes video and audio data to recognize the user's emotional state. Specifically, it identifies emotions from audio data using the Google® Cloud Speech-to-Text API and integrates them with facial expression data. The analysis results are used to provide more accurate learning support.

[0275] The server generates optimized feedback for the user based on the acquired emotional data and learning data. Specifically, it uses a generative AI model (e.g., GPT-4 (registered trademark)) to generate feedback such as advice, supplementary materials, additional problems, and words of encouragement. For example, if a user feels stressed while working on a math problem, it generates a message recommending "how to approach the problem calmly" or "taking a break."

[0276] An example of a prompt to input to a generative AI model is as follows:

[0277] "You recognize that the user is getting fatigued while solving a math problem. Generate feedback that includes the best relaxation techniques for him / her and some encouraging words."

[0278] The device receives the generated feedback from the server and displays it on the user's tablet screen. Specific study advice and the next task to tackle based on the analysis results and emotional state are displayed on the screen. The user can review this and proceed with their studies based on it. For example, a message might be displayed saying, "This part is difficult, but try to stay calm. It's also recommended that you take a break."

[0279] As described above, the system of the present invention captures and analyzes learners' learning activities in real time and provides feedback that takes into account their emotional state, thereby realizing high-quality, individually optimized learning support.

[0280] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0281] Step 1:

[0282] When a user begins learning, a camera installed on the desk automatically starts operating. The camera captures the user's learning activities in real time and constantly acquires video data. The input is the user's learning activities, and the output is video data. This data is used in the next step.

[0283] Step 2:

[0284] A camera connected to a terminal receives captured video data. The received video data is compressed and sent to a server using a secure communication protocol (e.g., HTTPS). The input is video data from the camera, and the output is compressed video data sent to the server. This process minimizes data communication delays.

[0285] Step 3:

[0286] The server analyzes the received video data using a specific software algorithm (e.g., OpenCV). The input here is the received video data, and the output is an analysis result showing changes in the user's facial expressions, movements, and gaze. Specifically, facial recognition and gaze tracking within the video data are performed. Based on this analysis, the user's concentration and fatigue levels are evaluated.

[0287] Step 4:

[0288] The server analyzes the user's emotional state using emotion recognition. It receives video and audio data as input, and obtains emotional data as output. Specifically, it identifies emotions from audio data using the Google Cloud Speech-to-Text API and integrates this with facial expression data. For example, it can read stress or excitement from the tone and tempo of the user's voice.

[0289] Step 5:

[0290] The server generates feedback based on the acquired emotion data and analysis data. It takes the analysis results and emotion data as input, and generates feedback as output using a generative AI model (e.g., GPT-4). An example of a prompt sentence to input to the generative AI model is as follows:

[0291] "You recognize that the user is getting fatigued while solving a math problem. Generate feedback that includes the best relaxation techniques for him / her and some encouraging words."

[0292] This step generates feedback that may include specific advice, supplementary materials, additional questions, or words of encouragement.

[0293] Step 6:

[0294] The device receives the generated feedback sent from the server and displays it on the user's tablet screen. The input is feedback data from the server, and the output is specific study advice and the next task the user should tackle. The user can then proceed with their studies based on this. For example, a message such as "This part is difficult, but try to stay calm. It is also recommended that you take a break" may be displayed.

[0295] (Application example 2)

[0296] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0297] Conventional learning support systems focus on assessing learners' levels of concentration and fatigue, but they also need to provide appropriate feedback in real time, not only to learners but also in other environments (e.g., customer service in brick-and-mortar stores). Such systems require advanced analytical techniques that take into account human emotions and behavior, which has been difficult to achieve with conventional technologies. It has also been difficult to develop a system that can accurately grasp a customer's emotional state, such as interest or confusion, in a store and provide appropriate advice on the spot.

[0298] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for capturing and evaluating the behavior and facial expressions of customers in the store, means including a generative AI model for generating optimal advice for the customer based on the evaluation results, and means for evaluating the customer's interest and confusion and generating appropriate prompt sentences. This makes it possible to analyze the customer's emotions and behavior in real time and provide appropriate advice on the spot.

[0299] A "learner" is an individual who uses a learning support system to carry out learning activities.

[0300] A "customer" is an individual who expresses interest in a product or service within a physical store.

[0301] "Photography means" refers to cameras or smart glasses used to capture the behavior and expressions of learners and customers.

[0302] The "transmission means" is a communication device for transmitting the video data captured by the image capture means to the server.

[0303] The "analysis means" refers to software and algorithms that analyze the video data received by the server and evaluate the behavior and emotions of learners and customers.

[0304] "Generation means" refers to a generative AI model that generates feedback appropriate for learners and customers based on the analysis results.

[0305] The "delivery means" is a display or notification system for providing the generated feedback to learners or customers.

[0306] A "generative AI model" is an artificial intelligence model that automatically generates optimal feedback and advice based on analysis results.

[0307] A "prompt" is a sentence that is input into a generative AI model and is an instruction sentence used to generate feedback or advice.

[0308] This invention is a system that can be applied not only to learners' learning activities but also to customer service support in brick-and-mortar stores. The system of the present invention is characterized by analyzing the behavior, facial expressions, and emotional states of learners and customers in real time and providing appropriate feedback and advice.

[0309] Hardware and Software

[0310] The main hardware used in this system is as follows:

[0311] Capture method: camera or smart glasses (e.g., Google Glass (registered trademark))

[0312] Transmission method: Tablet device or smartphone

[0313] Server: High-performance server (e.g., AWS EC2 instance)

[0314] The software used includes:

[0315] Video data analysis software: OpenCV

[0316] Sentiment analysis engine: Microsoft® Azure® Emotion API

[0317] Generative AI model: OpenAI's GPT-3 model

[0318] Data processing and calculation

[0319] The server performs the following steps:

[0320] 1. Acquisition of video data: Video data acquired by a camera is sent to a terminal. This data includes the behavior and facial expressions of the learner or customer.

[0321] 2. Data transmission: The device transmits the captured video data to the server, using data compression and a secure communication protocol.

[0322] 3. Video analysis and emotion assessment: The server analyzes the video data using OpenCV and assesses the emotional state using the Microsoft Azure Emotion API. From the analysis results, it identifies the learner's concentration level or fatigue level, and the customer's interest or confusion.

[0323] 4. Feedback generation: Based on the analysis results, the generative AI model (GPT-3) generates specific feedback and advice, including study advice and product information.

[0324] 5. Providing advice: The generated feedback is displayed and notified on the tablet device or smart glasses that are used as the means of providing the advice.

[0325] Examples and prompts

[0326] For example, in a brick-and-mortar scenario, consider a situation where a customer is interested in a particular product but is confused by its features. The server can analyze this confusion from the customer's facial expressions and behavior, and then use a generative AI model to input the following prompt:

[0327] "A customer was observed on camera. They were looking at product X with interest, but seemed confused about its features. Please generate additional information to provide to the customer."

[0328] Here, the generative AI model generates specific product descriptions and advice, which are displayed on the smart glasses, allowing store staff to provide information to customers at the right time.

[0329] In this way, a system can be built that provides optimal feedback to customers and learners in real time, achieving high-quality service and learning support.

[0330] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0331] Step 1:

[0332] The user begins learning activities or in-store activities. The imaging means (camera or smart glasses) captures the user's actions and facial expressions in real time. The input is the user's learning activities or in-store activities and facial expressions, and the output is video data.

[0333] Step 2:

[0334] The captured video data is sent to a device (tablet or smartphone). The device compresses the video data and sends it to a server using a secure communication protocol. The input is uncompressed video data, and the output is compressed video data.

[0335] Step 3:

[0336] The server analyzes the received video data. It processes visual information using OpenCV and evaluates emotional states using the Microsoft Azure Emotion API. From the analysis results, it identifies the user's behavior, facial expressions, concentration level, fatigue level, and customer interest or confusion. The input is compressed video data, and the output is analyzed behavioral and emotional data.

[0337] Step 4:

[0338] Based on the analysis results, the server uses a generative AI model (GPT-3) to generate specific feedback and advice. The prompt sentence is input into the generative AI model to obtain feedback appropriate for the specific situation. The input is the analyzed behavioral data, emotional data, and the prompt sentence, and the output is specific feedback and advice.

[0339] Step 5:

[0340] The means of providing this feedback (tablet device or smart glasses) displays the generated feedback and advice to the user, who then modifies or continues their learning or behavior accordingly. The input is the generated feedback and advice, and the output is the feedback provided to the user.

[0341] This process allows users to receive appropriate feedback in real time, enabling efficient learning and a meaningful purchasing experience.

[0342] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0343] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0344] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0345] [Second embodiment]

[0346] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0347] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0348] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0349] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0350] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0351] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0352] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0353] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0354] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0355] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0356] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0357] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0358] The present invention is a system for supporting learners' learning, specifically, a system that analyzes and supports learners' activities in real time using cameras and tablet devices installed in the learning environment. An embodiment of this system will be described in detail.

[0359] This system mainly comprises the following components: an imaging means, a transmitting means, an analyzing means, a generating means, and a providing means.

[0360] Filming method

[0361] Camera: A camera installed on each student's desk constantly records their learning activities. The camera is connected to a tablet device via Wi-Fi and captures video data in real time.

[0362] Transmission method

[0363] Device: The tablet device receives the video data captured by the camera, compresses it, and sends it to the server using a secure communication protocol. The video data is high resolution and can capture the learner's detailed movements and facial expressions.

[0364] Analysis means

[0365] Server: The server analyzes the received video data. Specific software algorithms are used to analyze the learner's facial expressions, movements, and changes in gaze. This analysis makes it possible to assess the learner's level of concentration and fatigue. Furthermore, by integrating and analyzing past learning data (class notes, past questions, mock test results, etc.), the server can identify the learner's strengths and weaknesses.

[0366] generation means

[0367] Server: Based on the analysis results, it generates optimized feedback for the learner. A generative AI model is used in this process to generate specific advice, supplementary materials, additional problems, etc. For example, if a learner is having trouble solving an equation, an explanation or practice problem will be provided to encourage review of that part.

[0368] Providing means

[0369] Device: The generated feedback is sent to a tablet device and displayed to the user. Specific study advice and the next task to tackle based on the analysis results are displayed on the tablet screen. The user can check this and continue studying.

[0370] Specific examples

[0371] For example, consider a user learning how to solve mathematical equations. A camera captures the user's learning, and the video data is sent to a server via a tablet device. The server analyzes the video data and detects that the user is making a calculation error in a particular step. It also identifies that the user has repeatedly made mistakes in this area based on past practice test results.

[0372] Based on this, the server generates detailed explanations for the user on "typical mistakes in solving equations and how to correct them" and sends them to the tablet device. The user can then review the explanations on the tablet and work on practice problems to correct the mistakes. This allows the user to receive specific, individually optimized learning support in real time.

[0373] As described above, the system of the present invention realizes high-quality learning support by capturing and analyzing learners' learning activities in real time and providing individually optimized feedback.

[0374] The processing flow will be explained below.

[0375] Step 1:

[0376] User: Starts the learning application on the tablet device and configures the camera connection settings.

[0377] Step 2:

[0378] Device: The tablet device connects to the tabletop camera via Wi-Fi and checks the camera's connection status.

[0379] Step 3:

[0380] Camera: The tabletop camera continuously captures the learner's learning activities, capturing images at a set frame rate and resolution.

[0381] Step 4:

[0382] Terminal: Compresses video data captured in real time and sends it to the server using a secure communication protocol.

[0383] Step 5:

[0384] Server: Receives compressed video data and temporarily stores it in a database.

[0385] Step 6:

[0386] Server: Using "Gemini," the received video data is analyzed. Specifically, changes in the learner's facial expressions, movements, and gaze are analyzed to evaluate the learner's level of concentration and fatigue.

[0387] Step 7:

[0388] Server: Based on the analysis results, the learner's strengths and weaknesses are identified, and past learning data (class notes, past questions, mock test results, etc.) is integrated as necessary for analysis.

[0389] Step 8:

[0390] Server: Based on the analysis results, a generative AI model generates optimized feedback for the learner, including specific advice, supplementary materials, and additional questions.

[0391] Step 9:

[0392] Server: Sends the generated feedback to the tablet device.

[0393] Step 10:

[0394] Device: The tablet device receives the feedback and displays it in the application's user interface.

[0395] Step 11:

[0396] User: Check the feedback displayed on the tablet and learn what to study next or what explanations to give.

[0397] Step 12:

[0398] Server: Records new progress data of learners and stores it in a database. Based on the progress, the learning plan and feedback are updated.

[0399] Step 13:

[0400] User: Check the learning progress and continue learning according to the plan. If necessary, repeat the cycle from step 1.

[0401] Example 1

[0402] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0403] Conventional learning support systems have difficulty analyzing learners' learning activities in detail in real time, making it difficult to provide individually optimized feedback in a timely manner. Furthermore, they lack the functionality to capture subtle changes in learners' facial expressions and gaze to evaluate their concentration and fatigue levels, making it difficult to optimally improve learners' performance.

[0404] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0405] In this invention, the server includes a photographing means, a transmitting means, an analyzing means, a generating means, and a providing means. This enables the learner's learning activities to be analyzed in real time and individually optimized feedback to be provided promptly. Specifically, the server uses a specific software algorithm to analyze the received video data and analyzes the learner's facial expressions, movements, and changes in gaze to evaluate the learner's level of concentration and fatigue. Based on the results, the server uses a generative AI model to generate feedback appropriate for the learner, and has a means for transmitting and displaying the feedback on the mobile information terminal. This allows the learner to receive optimal learning support in real time, improving the quality of their learning.

[0406] The "filming means" is a device that films the learning activities of the learner in real time and acquires video data.

[0407] The "transmission means" is a device or software that has the function of compressing the video data acquired by the image capture means and transmitting it to the server using a secure communication protocol.

[0408] "Analysis means" refers to a device or software that has the function of analyzing the video data received by the server using a specific software algorithm, and analyzing changes in the learner's facial expressions, movements, and gaze to evaluate their level of concentration and fatigue.

[0409] "Generation means" refers to a device or software that has the function of utilizing a generative AI model to automatically generate feedback optimized for the learner based on the analysis results obtained by the analysis means and past learning data.

[0410] The "providing means" is a device or software that has the function of transmitting the feedback generated by the generating means to the learner's mobile information terminal and displaying it.

[0411] A "generative AI model" is a model that uses artificial intelligence algorithms to generate appropriate feedback based on input data.

[0412] This system is constructed to analyze learners' learning activities in detail in real time and provide individually optimized feedback. The system consists of a capturing means, a transmitting means, an analyzing means, a generating means, and a providing means.

[0413] Filming method

[0414] Camera: A camera (e.g., a high-resolution webcam) installed on the student's desk continuously records the student's learning activities. The camera is connected to a tablet device via Wi-Fi and captures video data in real time.

[0415] Transmission method

[0416] Terminal: The tablet terminal receives the video data captured by the camera and automatically compresses it using an efficient video compression algorithm such as H.264. The compressed video data is then sent to the server using a secure communication protocol (e.g., HTTPS).

[0417] Analysis means

[0418] Server: The server analyzes the received video data. For example, image processing software algorithms such as OpenCV are used for the analysis. The analysis method detects changes in the learner's facial expressions, movements, and line of sight, and evaluates the learner's level of concentration and fatigue based on these.

[0419] generation means

[0420] Server: Based on the analysis results and past learning data, a generative AI model (e.g., a generative AI model) is used to generate optimized feedback for the learner. For example, if a learner is having trouble solving an equation, an explanation or exercise tailored to the content is generated.

[0421] Providing means

[0422] Device: The generated feedback is sent to the learner's tablet device and displayed to them. Specific study advice and the next task to be tackled based on the analysis results are displayed on the tablet screen.

[0423] Specific examples

[0424] For example, if a user is learning how to solve mathematical equations, a camera captures the user's learning activities and the video data is sent to a server via a tablet device. The server analyzes the video data and detects that the user is making a calculation error in a specific step. It also identifies that the user has repeatedly made mistakes in this area based on past practice test results. Based on this, the server generates detailed explanations for the user on "typical mistakes in solving equations and how to correct them" and sends them to the tablet device. The user then checks the explanations on the tablet and works on practice problems to correct their mistakes.

[0425] Prompt Sentence Examples

[0426] "Analyze where a user makes mistakes in solving math equations and generate specific explanations and practice questions for review."

[0427] "Calculate the learner's concentration and fatigue levels from video data and provide appropriate learning advice."

[0428] The above is an embodiment of the present invention. This system enables learners to receive optimal learning support in real time, thereby improving the quality of their learning.

[0429] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0430] Step 1:

[0431] Filming learning activities

[0432] Camera: A camera installed on each student's desk constantly records their learning activities, capturing in high resolution the student's facial expressions, movements, and gaze in detail.

[0433] Input: Real-time video of the learner

[0434] Output: High-resolution video data of learning activities

[0435] Step 2:

[0436] Video data compression and transmission

[0437] Terminal: The tablet terminal compresses the video data received from the camera using the H.264 algorithm, and then transmits the compressed data to the server using a secure communication protocol (e.g., HTTPS).

[0438] Input: High-resolution video of learning activities

[0439] Output: Compressed video data

[0440] Step 3:

[0441] Video data analysis

[0442] Server: Analyzes the received compressed video data. Image processing algorithms such as OpenCV are used for the analysis. The server analyzes the learner's facial expressions, movements, and changes in gaze to evaluate the learner's level of concentration and fatigue. It also compares this with past learning data.

[0443] Input: Compressed video data, past learning data

[0444] Output: Evaluation results of learner's concentration, fatigue, strengths and weaknesses

[0445] Step 4:

[0446] Generate feedback

[0447] Server: Based on the analysis results, the server sends a prompt to the generative AI model. The prompt includes the learner's situation and the feedback they need. The generative AI model generates specific feedback (explanations, practice questions, advice, etc.) based on the prompt.

[0448] Input: Analysis results, prompt text

[0449] Output: Generated feedback

[0450] Step 5:

[0451] Providing Feedback

[0452] Terminal: The feedback sent from the server is displayed on the tablet terminal. The user checks the feedback displayed on the tablet. Examples include "typical mistakes in solving equations and how to correct them."

[0453] Input: Generated feedback

[0454] Output: Feedback display on tablet device

[0455] Examples:

[0456] As an example, consider the case where a user is learning how to solve a mathematical equation.

[0457] Step 1: The camera captures the user's learning activity and captures it as high-resolution video data.

[0458] Step 2: The terminal compresses the video data using the H.264 algorithm and sends it to the server using the HTTPS protocol.

[0459] Step 3: The server analyzes the received video data using OpenCV to detect if the user is making an incorrect calculation at a specific step. It also integrates past learning data and evaluates the learner's strengths and weaknesses based on this.

[0460] Step 4: Based on the analysis results, the generative AI model is prompted to "analyze the user's mistakes in solving mathematical equations and generate specific explanations and practice questions for review." The generative AI model then generates specific feedback.

[0461] Step 5: The generated feedback is sent to the device, and "typical mistakes in solving equations and how to correct them" are displayed on the user's tablet. The user then tries learning again while looking at the results.

[0462] The above are the specific processing steps of the system of the present invention, which enable learners to receive optimal learning support in real time.

[0463] (Application example 1)

[0464] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0465] Conventional automated equipment operating in factories lacks real-time analysis functions to optimize efficiency and precision, and is unable to provide immediate feedback on improvements. As a result, waste and malfunctions tend to occur in the equipment's operation, leading to problems such as reduced productivity and poor quality.

[0466] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0467] In this invention, the server includes a photographing means for photographing the learning activities of the learner, a transmitting means for transmitting video data photographed by the photographing means to the server, an analyzing means for analyzing the received video data, a generating means for generating feedback appropriate for the learner based on the analysis results, a providing means for providing the generated feedback to the learner, an automation device operation analyzing means for analyzing the operation of automation devices operating in the factory and providing feedback on areas for improvement in efficiency and accuracy, and a display means for displaying the feedback. This makes it possible to monitor the operation of automation devices operating in the factory in real time and improve efficiency and accuracy based on the analysis results.

[0468] "Photographing means" refers to a camera or other image capturing device that captures the learning activities of learners and the operation of automated equipment in a factory.

[0469] The "transmission means" refers to a device or technology for transmitting the video data captured by the image capture means to the server, and is, for example, a device that uses a tablet terminal or a communication protocol.

[0470] The "analysis means" is a technology for analyzing the video data received in the server and evaluating the activities of the learners and the operation status of the automated devices.

[0471] "Generation means" refers to technology for generating appropriate feedback to learners and improvements to automated devices based on the analysis results.

[0472] "Providing means" refers to the devices and technologies used to display the generated feedback to the learner or operator.

[0473] "Automated equipment operation analysis means" is a technology for analyzing the operation of automated equipment operating in a factory and detecting areas for improvement in efficiency and accuracy.

[0474] "Display means" refers to a device or technology for visually presenting the analysis and generated feedback to the user, such as a tablet device or display.

[0475] "Real-time" refers to a time frame in which data is acquired, transmitted, analyzed, and feedback generated nearly simultaneously.

[0476] "Efficiency" is an index that indicates the degree to which automated equipment operating in a factory operates optimally and without waste.

[0477] "Accuracy" is a measure of the ability or degree to which an automated device accurately performs a required operation.

[0478] The present invention is a system for optimizing automated equipment in a factory by analyzing its operation and providing feedback on areas for improvement in efficiency and accuracy. This system is realized through the following steps.

[0479] System Program

[0480] This system mainly comprises an imaging means, a transmission means, an analysis means, a generation means, and a provision means.

[0481] Filming method

[0482] A camera is used as a recording medium. The camera records the operation of the automated equipment in real time and acquires the video data. The camera is connected to a tablet device via Wi-Fi or a wired connection.

[0483] Transmission method

[0484] The device receives the captured video data, compresses it, and transmits it to a server using a secure communication protocol. The high-resolution video data can capture even the smallest details of the automated equipment's operation.

[0485] Analysis means

[0486] The server analyzes the received video data using specific software algorithms (e.g., computer vision technology) to evaluate the efficiency and accuracy of the automated equipment's operation. This analysis identifies waste and malfunctions in the equipment's operation.

[0487] generation means

[0488] Based on the analysis results, feedback is generated to improve efficiency and accuracy. A generative AI model is used in this process to generate specific improvements and supplemental materials. For example, feedback such as "The movement of the right arm is slow and needs adjustment" is generated.

[0489] Providing means

[0490] The generated feedback is sent to the tablet device and displayed to the user. The feedback is provided in a visual format and includes specific instructions on how to improve.

[0491] Hardware and software used

[0492] Camera: Films the operation of automated equipment.

[0493] Tablet device: Receives video data captured by the camera and sends it to the server.

[0494] Server: Analyzes video data and generates feedback using computer vision techniques and generative AI models.

[0495] Specific examples

[0496] For example, if there is an automated machine that assembles a product, a camera will record its operation. The server will analyze the video data and identify the problem: "The right arm is moving slowly." Based on the analysis results, the generative AI model will generate feedback such as "Please adjust the right arm to move 10% faster," and send this to a tablet device. The user will then check this feedback on the tablet and adjust the machine's operation.

[0497] Prompt Sentence Examples

[0498] Configure your robot to analyze footage of its movements and identify areas for improvement in efficiency and accuracy. Analyze the movements in the footage and generate feedback that is user-friendly and specific.

[0499] This completes the detailed description of the embodiment of the present invention. This system makes it possible to monitor the operation of automated devices operating in a factory in real time and improve efficiency and accuracy based on the analysis results.

[0500] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0501] Step 1:

[0502] Camera image capture

[0503] Input: Factory automation equipment operation

[0504] Output: High-resolution video data

[0505] The camera captures the operation of the automated equipment in real time, capturing the operating status of the automated equipment and sending the video to a terminal as high-resolution video data.

[0506] Step 2:

[0507] Receive video data on the device

[0508] Input: High-resolution video data from a camera

[0509] Output: Compressed video data stored on the device

[0510] The device receives the video data captured by the camera and compresses it, making it ready to be sent to the server.

[0511] Step 3:

[0512] Compressed video data sent to server

[0513] Input: Compressed video data

[0514] Output: Video data sent to the server

[0515] The device transmits the compressed video data to the server using a secure communication protocol, with encryption techniques used to ensure data integrity and privacy.

[0516] Step 4:

[0517] Video data analysis by server

[0518] Input: Received video data

[0519] Output: Evaluation of the efficiency and accuracy of the operation as a result of the analysis

[0520] The server analyzes the received video data using computer vision techniques and performs calculations on the data to evaluate the operational status of the automated device (e.g., speed, accuracy of operation).

[0521] Step 5:

[0522] Feedback Generation

[0523] Input: Analysis results

[0524] Output: Feedback on areas for improvement in efficiency and accuracy

[0525] The server generates feedback based on the analysis results to improve efficiency and accuracy, using a generative AI model to generate specific improvements and supplementary materials.

[0526] Step 6:

[0527] Provide feedback to the device

[0528] Input: Generated feedback

[0529] Output: Feedback displayed on the terminal

[0530] The server sends the generated feedback to the device, which receives it and visually displays it to the user, allowing the user to confirm and implement specific improvements.

[0531] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0532] The present invention is a system for supporting learners in their learning, and is characterized by its ability to provide feedback based on the learner's emotional state by incorporating an emotion engine. An embodiment of this system will be described in detail.

[0533] This system mainly consists of the following components: a photographing means, a transmitting means, an analyzing means, a generating means, a providing means, and an emotion engine.

[0534] Filming method

[0535] Camera: A camera installed on each student's desk constantly records their learning activities. The camera is connected to a tablet device via Wi-Fi and captures video data in real time.

[0536] Transmission method

[0537] Terminal: The tablet terminal receives the video data captured by the camera, compresses it, and sends it to the server using a secure communication protocol.

[0538] Analysis means

[0539] Server: The server analyzes the received video data. Specific software algorithms are used to analyze the learner's facial expressions, movements, and changes in gaze. This analysis evaluates the learner's level of concentration and fatigue. It also integrates and analyzes past learning data (class notes, past questions, mock test results, etc.) to identify the learner's strengths and weaknesses.

[0540] Emotion Engine

[0541] Server: The emotion engine analyzes video and audio data to recognize the learner's emotional state. Specifically, it identifies the learner's emotions, such as joy, sadness, and stress, from facial expression and audio data. This emotional data is integrated with data obtained by the analysis means, enabling more accurate learning support.

[0542] generation means

[0543] Server: Generates optimized feedback for learners based on the analysis results and emotional data obtained by the emotion engine. A generative AI model is used in this process to generate specific advice, supplementary materials, additional questions, etc. For example, if a learner is feeling stressed, encouraging words and advice on how to relax are provided.

[0544] Providing means

[0545] Device: The generated feedback is sent to a tablet device and displayed to the user. Specific study advice and next tasks based on the analysis results and emotional state are displayed on the tablet screen. The user can confirm this and continue studying.

[0546] Specific examples

[0547] For example, consider a user learning how to solve a mathematical equation. A camera captures the user's learning, and the video data is sent to a server via a tablet device. The server analyzes the video data and detects that the user is making a calculation error at a specific step. Furthermore, the emotion engine recognizes from the user's facial expressions and voice that the user is feeling frustrated because they cannot solve the equation.

[0548] Based on this, the server generates detailed explanations for the user about "common mistakes in solving equations and how to correct them." It also adds encouraging words to the explanations, such as "This part is difficult, but try to stay calm. It's also recommended to take breaks." The generated feedback is sent to a tablet device, where the user can review it and work on practice problems to correct their mistakes. This allows the user to receive specific, emotionally appropriate feedback in real time.

[0549] As described above, the system of the present invention captures and analyzes learners' learning activities in real time and provides feedback that takes into account their emotional state, thereby realizing high-quality, individually optimized learning support.

[0550] The processing flow will be explained below.

[0551] Step 1:

[0552] User: Starts the learning application on the tablet device and configures the camera connection settings.

[0553] Step 2:

[0554] Device: The tablet device connects to the tabletop camera via Wi-Fi and checks the camera's connection status.

[0555] Step 3:

[0556] Camera: The tabletop camera continuously captures the learner's learning activities, capturing images at a set frame rate and resolution.

[0557] Step 4:

[0558] Terminal: Compresses video data captured in real time and sends it to the server using a secure communication protocol.

[0559] Step 5:

[0560] Server: Receives compressed video data and temporarily stores it in a database.

[0561] Step 6:

[0562] Server: Analyzes the received video data using the "analysis means." Specifically, it analyzes the learner's facial expressions, movements, and changes in line of sight to evaluate the learner's level of concentration and fatigue.

[0563] Step 7:

[0564] Server: Obtains past learning data (class notes, past questions, mock test results, etc.) from a database and integrates it with current analysis results to identify the learner's strengths and weaknesses.

[0565] Step 8:

[0566] Server: Using an "emotion engine," it analyzes video and audio data to recognize the learner's emotional state (happiness, sadness, stress, etc.). Specifically, it runs an algorithm that identifies emotions from facial expression and audio data.

[0567] Step 9:

[0568] Server: The "generator" generates optimal feedback based on the analysis results and emotional data. For example, if a calculation error is detected, it will include detailed instructions on how to correct the error and advice corresponding to the user's emotional state.

[0569] Step 10:

[0570] Server: Sends the generated feedback to the tablet device.

[0571] Step 11:

[0572] Device: The tablet device receives the feedback and displays it in the application's user interface.

[0573] Step 12:

[0574] User: Check the feedback displayed on the tablet and learn what to study next or what explanations to give. If necessary, follow the advice to relax.

[0575] Step 13:

[0576] Server: Records new progress and emotion data of the learner and stores it in a database. Based on the progress and emotion data, the learning plan and feedback content are updated.

[0577] Step 14:

[0578] User: Check their learning progress and emotional state, continue learning according to the plan, and repeat the cycle from step 1 if necessary.

[0579] Example 2

[0580] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0581] Conventional learning support systems can record learners' learning activities and provide feedback, but they have limitations in providing real-time support that takes into account the learner's emotional state. This makes it difficult to provide feedback appropriate to a learner's emotional state when they feel stressed or fatigued. Furthermore, there is a need for systems that go beyond simple data analysis to automatically identify a learner's strengths and weaknesses and provide individually optimized learning support based on that information.

[0582] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0583] In this invention, the server includes a photographing means for photographing the learning activities of the learner, a transmitting means for transmitting the video data photographed by the photographing means to the server, an analyzing means in the server for analyzing the received video data, a generating means for generating appropriate feedback based on the data and emotional data obtained by the analyzing means, a providing means for providing the generated feedback to the learner, an emotion recognition means for analyzing the emotional state of the learner from the video data and audio data, and a generating means using a generative AI model for integrating the obtained emotional data and data from the analyzing means to generate feedback optimized for the learner. This enables high-quality learning support that is individually optimized in real time and takes into account the emotional state of the learner.

[0584] "Photography means" refers to devices used to record learners' learning activities, such as cameras.

[0585] The "transmission means" refers to a device or function for transmitting the video data captured by the image capture means to the server.

[0586] "Analysis means" refers to software algorithms or programs used to analyze received video data and analyse changes in the learner's facial expressions, movements and gaze.

[0587] The "generation means" refers to a process or device that generates optimized feedback for the learner based on the data and emotional data obtained by the analysis means.

[0588] "Providing means" refers to a device or method for presenting the generated feedback to the learner.

[0589] "Emotion recognition means" refers to software or algorithms that analyze video and audio data to identify a learner's emotional state.

[0590] "Generative AI model" refers to an artificial intelligence model used to generate optimal feedback and advice for learners based on collected data.

[0591] "Feedback" refers to specific advice, supplementary materials, additional questions, words of encouragement, etc., generated by the analysis means or generation means and provided to the learner.

[0592] "Learning progress management means" refers to a process or device for recording a learner's progress and adjusting the learning plan based on the saved progress data.

[0593] The present invention is a system that comprehensively supports learners' learning activities, and its distinctive feature is that it incorporates the learner's emotional state. A specific embodiment of this system will be described. The main components include an image capture means, a transmission means, an analysis means, a generation means, a provision means, and an emotion recognition means. This makes it possible to provide learners with individually optimized feedback in real time.

[0594] First, when a user begins studying, a camera installed on the desk automatically starts operating. The camera captures the user's study activities in real time and continuously acquires video data. The camera is connected to a tablet device via Wi-Fi, and the captured video data is sent to the tablet device.

[0595] A camera connected to a terminal sends captured video data to the terminal. The terminal compresses the received video data and sends it to a server via a secure communication protocol (e.g., HTTPS). Data compression minimizes communication delays.

[0596] The server analyzes the received video data. Specifically, it uses OpenCV and other tools to analyze the user's facial expressions, movements, and changes in gaze. This allows it to evaluate the user's level of concentration and fatigue. It also integrates and analyzes past learning data (class notes, past questions, mock test results, etc.) to identify the user's strengths and weaknesses.

[0597] The emotion recognition means analyzes video and audio data to recognize the user's emotional state. Specifically, it identifies emotions from audio data using the Google Cloud Speech-to-Text API and integrates them with facial expression data. The results of this analysis are used to provide more accurate learning support.

[0598] The server generates optimized feedback for the user based on the acquired emotional data and learning data. Specifically, it uses a generative AI model (e.g., GPT-4) to generate feedback such as advice, supplementary materials, additional problems, and words of encouragement. For example, if a user feels stressed while working on a math problem, it generates a message recommending "how to approach the problem calmly" or "taking a break."

[0599] An example of a prompt to input to a generative AI model is as follows:

[0600] "You recognize that the user is getting fatigued while solving a math problem. Generate feedback that includes the best relaxation techniques for him / her and some encouraging words."

[0601] The device receives the generated feedback from the server and displays it on the user's tablet screen. Specific study advice and the next task to tackle based on the analysis results and emotional state are displayed on the screen. The user can review this and proceed with their studies based on it. For example, a message might be displayed saying, "This part is difficult, but try to stay calm. It's also recommended that you take a break."

[0602] As described above, the system of the present invention captures and analyzes learners' learning activities in real time and provides feedback that takes into account their emotional state, thereby realizing high-quality, individually optimized learning support.

[0603] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0604] Step 1:

[0605] When a user begins learning, a camera installed on the desk automatically starts operating. The camera captures the user's learning activities in real time and constantly acquires video data. The input is the user's learning activities, and the output is video data. This data is used in the next step.

[0606] Step 2:

[0607] A camera connected to a terminal receives captured video data. The received video data is compressed and sent to a server using a secure communication protocol (e.g., HTTPS). The input is video data from the camera, and the output is compressed video data sent to the server. This process minimizes data communication delays.

[0608] Step 3:

[0609] The server analyzes the received video data using a specific software algorithm (e.g., OpenCV). The input here is the received video data, and the output is an analysis result showing changes in the user's facial expressions, movements, and gaze. Specifically, facial recognition and gaze tracking within the video data are performed. Based on this analysis, the user's concentration and fatigue levels are evaluated.

[0610] Step 4:

[0611] The server analyzes the user's emotional state using emotion recognition. It receives video and audio data as input, and obtains emotional data as output. Specifically, it identifies emotions from audio data using the Google Cloud Speech-to-Text API and integrates this with facial expression data. For example, it can read stress or excitement from the tone and tempo of the user's voice.

[0612] Step 5:

[0613] The server generates feedback based on the acquired emotion data and analysis data. It takes the analysis results and emotion data as input, and generates feedback as output using a generative AI model (e.g., GPT-4). An example of a prompt sentence to input to the generative AI model is as follows:

[0614] "You recognize that the user is getting fatigued while solving a math problem. Generate feedback that includes the best relaxation techniques for him / her and some encouraging words."

[0615] This step generates feedback that may include specific advice, supplementary materials, additional questions, or words of encouragement.

[0616] Step 6:

[0617] The device receives the generated feedback sent from the server and displays it on the user's tablet screen. The input is feedback data from the server, and the output is specific study advice and the next task the user should tackle. The user can then proceed with their studies based on this. For example, a message such as "This part is difficult, but try to stay calm. It is also recommended that you take a break" may be displayed.

[0618] (Application example 2)

[0619] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0620] Conventional learning support systems focus on assessing learners' levels of concentration and fatigue, but they also need to provide appropriate feedback in real time, not only to learners but also in other environments (e.g., customer service in brick-and-mortar stores). Such systems require advanced analytical techniques that take into account human emotions and behavior, which has been difficult to achieve with conventional technologies. It has also been difficult to develop a system that can accurately grasp a customer's emotional state, such as interest or confusion, in a store and provide appropriate advice on the spot.

[0621] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for capturing and evaluating the behavior and facial expressions of customers in the store, means including a generative AI model for generating optimal advice for the customer based on the evaluation results, and means for evaluating the customer's interest and confusion and generating appropriate prompt sentences. This makes it possible to analyze the customer's emotions and behavior in real time and provide appropriate advice on the spot.

[0622] A "learner" is an individual who uses a learning support system to carry out learning activities.

[0623] A "customer" is an individual who expresses interest in a product or service within a physical store.

[0624] "Photography means" refers to cameras or smart glasses used to capture the behavior and expressions of learners and customers.

[0625] The "transmission means" is a communication device for transmitting the video data captured by the image capture means to the server.

[0626] The "analysis means" refers to software and algorithms that analyze the video data received by the server and evaluate the behavior and emotions of learners and customers.

[0627] "Generation means" refers to a generative AI model that generates feedback appropriate for learners and customers based on the analysis results.

[0628] The "delivery means" is a display or notification system for providing the generated feedback to learners or customers.

[0629] A "generative AI model" is an artificial intelligence model that automatically generates optimal feedback and advice based on analysis results.

[0630] A "prompt" is a sentence that is input into a generative AI model and is an instruction sentence used to generate feedback or advice.

[0631] This invention is a system that can be applied not only to learners' learning activities but also to customer service support in brick-and-mortar stores. The system of the present invention is characterized by analyzing the behavior, facial expressions, and emotional states of learners and customers in real time and providing appropriate feedback and advice.

[0632] Hardware and Software

[0633] The main hardware used in this system is as follows:

[0634] Capture method: camera or smart glasses (e.g. Google Glass)

[0635] Transmission method: Tablet device or smartphone

[0636] Server: High-performance server (e.g. AWS EC2 instance)

[0637] The software used includes:

[0638] Video data analysis software: OpenCV

[0639] Sentiment analysis engine: Microsoft Azure Emotion API

[0640] Generative AI model: OpenAI's GPT-3 model

[0641] Data processing and calculation

[0642] The server performs the following steps:

[0643] 1. Acquisition of video data: Video data acquired by a camera is sent to a terminal. This data includes the behavior and facial expressions of the learner or customer.

[0644] 2. Data transmission: The device transmits the captured video data to the server, using data compression and a secure communication protocol.

[0645] 3. Video analysis and emotion assessment: The server analyzes the video data using OpenCV and assesses the emotional state using the Microsoft Azure Emotion API. From the analysis results, it identifies the learner's concentration level or fatigue level, and the customer's interest or confusion.

[0646] 4. Feedback generation: Based on the analysis results, the generative AI model (GPT-3) generates specific feedback and advice, including study advice and product information.

[0647] 5. Providing advice: The generated feedback is displayed and notified on the tablet device or smart glasses that are used as the means of providing the advice.

[0648] Examples and prompts

[0649] For example, in a brick-and-mortar scenario, consider a situation where a customer is interested in a particular product but is confused by its features. The server can analyze this confusion from the customer's facial expressions and behavior, and then use a generative AI model to input the following prompt:

[0650] "A customer was observed on camera. They were looking at product X with interest, but seemed confused about its features. Please generate additional information to provide to the customer."

[0651] Here, the generative AI model generates specific product descriptions and advice, which are displayed on the smart glasses, allowing store staff to provide information to customers at the right time.

[0652] In this way, a system can be built that provides optimal feedback to customers and learners in real time, achieving high-quality service and learning support.

[0653] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0654] Step 1:

[0655] The user begins learning activities or in-store activities. The imaging means (camera or smart glasses) captures the user's actions and facial expressions in real time. The input is the user's learning activities or in-store activities and facial expressions, and the output is video data.

[0656] Step 2:

[0657] The captured video data is sent to a device (tablet or smartphone). The device compresses the video data and sends it to a server using a secure communication protocol. The input is uncompressed video data, and the output is compressed video data.

[0658] Step 3:

[0659] The server analyzes the received video data. It processes visual information using OpenCV and evaluates emotional states using the Microsoft Azure Emotion API. From the analysis results, it identifies the user's behavior, facial expressions, concentration level, fatigue level, and customer interest or confusion. The input is compressed video data, and the output is analyzed behavioral and emotional data.

[0660] Step 4:

[0661] Based on the analysis results, the server uses a generative AI model (GPT-3) to generate specific feedback and advice. The prompt sentence is input into the generative AI model to obtain feedback appropriate for the specific situation. The input is the analyzed behavioral data, emotional data, and the prompt sentence, and the output is specific feedback and advice.

[0662] Step 5:

[0663] The means of providing this feedback (tablet device or smart glasses) displays the generated feedback and advice to the user, who then modifies or continues their learning or behavior accordingly. The input is the generated feedback and advice, and the output is the feedback provided to the user.

[0664] This process allows users to receive appropriate feedback in real time, enabling efficient learning and a meaningful purchasing experience.

[0665] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0666] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0667] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0668] [Third embodiment]

[0669] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0670] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0671] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0672] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0673] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0674] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0675] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0676] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0677] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0678] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0679] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0680] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0681] The present invention is a system for supporting learners' learning, specifically, a system that analyzes and supports learners' activities in real time using cameras and tablet devices installed in the learning environment. An embodiment of this system will be described in detail.

[0682] This system mainly comprises the following components: an imaging means, a transmitting means, an analyzing means, a generating means, and a providing means.

[0683] Filming method

[0684] Camera: A camera installed on each student's desk constantly records their learning activities. The camera is connected to a tablet device via Wi-Fi and captures video data in real time.

[0685] Transmission method

[0686] Device: The tablet device receives the video data captured by the camera, compresses it, and sends it to the server using a secure communication protocol. The video data is high resolution and can capture the learner's detailed movements and facial expressions.

[0687] Analysis means

[0688] Server: The server analyzes the received video data. Specific software algorithms are used to analyze the learner's facial expressions, movements, and changes in gaze. This analysis makes it possible to assess the learner's level of concentration and fatigue. Furthermore, by integrating and analyzing past learning data (class notes, past questions, mock test results, etc.), the server can identify the learner's strengths and weaknesses.

[0689] generation means

[0690] Server: Based on the analysis results, it generates optimized feedback for the learner. A generative AI model is used in this process to generate specific advice, supplementary materials, additional problems, etc. For example, if a learner is having trouble solving an equation, an explanation or practice problem will be provided to encourage review of that part.

[0691] Providing means

[0692] Device: The generated feedback is sent to a tablet device and displayed to the user. Specific study advice and the next task to tackle based on the analysis results are displayed on the tablet screen. The user can check this and continue studying.

[0693] Specific examples

[0694] For example, consider a user learning how to solve mathematical equations. A camera captures the user's learning, and the video data is sent to a server via a tablet device. The server analyzes the video data and detects that the user is making a calculation error in a particular step. It also identifies that the user has repeatedly made mistakes in this area based on past practice test results.

[0695] Based on this, the server generates detailed explanations for the user on "typical mistakes in solving equations and how to correct them" and sends them to the tablet device. The user can then review the explanations on the tablet and work on practice problems to correct the mistakes. This allows the user to receive specific, individually optimized learning support in real time.

[0696] As described above, the system of the present invention realizes high-quality learning support by capturing and analyzing learners' learning activities in real time and providing individually optimized feedback.

[0697] The processing flow will be explained below.

[0698] Step 1:

[0699] User: Starts the learning application on the tablet device and configures the camera connection settings.

[0700] Step 2:

[0701] Device: The tablet device connects to the tabletop camera via Wi-Fi and checks the camera's connection status.

[0702] Step 3:

[0703] Camera: The tabletop camera continuously captures the learner's learning activities, capturing images at a set frame rate and resolution.

[0704] Step 4:

[0705] Terminal: Compresses video data captured in real time and sends it to the server using a secure communication protocol.

[0706] Step 5:

[0707] Server: Receives compressed video data and temporarily stores it in a database.

[0708] Step 6:

[0709] Server: Using "Gemini," the received video data is analyzed. Specifically, changes in the learner's facial expressions, movements, and gaze are analyzed to evaluate the learner's level of concentration and fatigue.

[0710] Step 7:

[0711] Server: Based on the analysis results, the learner's strengths and weaknesses are identified, and past learning data (class notes, past questions, mock test results, etc.) is integrated as necessary for analysis.

[0712] Step 8:

[0713] Server: Based on the analysis results, a generative AI model generates optimized feedback for the learner, including specific advice, supplementary materials, and additional questions.

[0714] Step 9:

[0715] Server: Sends the generated feedback to the tablet device.

[0716] Step 10:

[0717] Device: The tablet device receives the feedback and displays it in the application's user interface.

[0718] Step 11:

[0719] User: Check the feedback displayed on the tablet and learn what to study next or what explanations to give.

[0720] Step 12:

[0721] Server: Records new progress data of learners and stores it in a database. Based on the progress, the learning plan and feedback are updated.

[0722] Step 13:

[0723] User: Check the learning progress and continue learning according to the plan. If necessary, repeat the cycle from step 1.

[0724] Example 1

[0725] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0726] Conventional learning support systems have difficulty analyzing learners' learning activities in detail in real time, making it difficult to provide individually optimized feedback in a timely manner. Furthermore, they lack the functionality to capture subtle changes in learners' facial expressions and gaze to evaluate their concentration and fatigue levels, making it difficult to optimally improve learners' performance.

[0727] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0728] In this invention, the server includes a photographing means, a transmitting means, an analyzing means, a generating means, and a providing means. This enables the learner's learning activities to be analyzed in real time and individually optimized feedback to be provided promptly. Specifically, the server uses a specific software algorithm to analyze the received video data and analyzes the learner's facial expressions, movements, and changes in gaze to evaluate the learner's level of concentration and fatigue. Based on the results, the server uses a generative AI model to generate feedback appropriate for the learner, and has a means for transmitting and displaying the feedback on the mobile information terminal. This allows the learner to receive optimal learning support in real time, improving the quality of their learning.

[0729] The "filming means" is a device that films the learning activities of the learner in real time and acquires video data.

[0730] The "transmission means" is a device or software that has the function of compressing the video data acquired by the image capture means and transmitting it to the server using a secure communication protocol.

[0731] "Analysis means" refers to a device or software that has the function of analyzing the video data received by the server using a specific software algorithm, and analyzing changes in the learner's facial expressions, movements, and gaze to evaluate their level of concentration and fatigue.

[0732] "Generation means" refers to a device or software that has the function of utilizing a generative AI model to automatically generate feedback optimized for the learner based on the analysis results obtained by the analysis means and past learning data.

[0733] The "providing means" is a device or software that has the function of transmitting the feedback generated by the generating means to the learner's mobile information terminal and displaying it.

[0734] A "generative AI model" is a model that uses artificial intelligence algorithms to generate appropriate feedback based on input data.

[0735] This system is constructed to analyze learners' learning activities in detail in real time and provide individually optimized feedback. The system consists of a capturing means, a transmitting means, an analyzing means, a generating means, and a providing means.

[0736] Filming method

[0737] Camera: A camera (e.g., a high-resolution webcam) installed on the student's desk continuously records the student's learning activities. The camera is connected to a tablet device via Wi-Fi and captures video data in real time.

[0738] Transmission method

[0739] Terminal: The tablet terminal receives the video data captured by the camera and automatically compresses it using an efficient video compression algorithm such as H.264. The compressed video data is then sent to the server using a secure communication protocol (e.g., HTTPS).

[0740] Analysis means

[0741] Server: The server analyzes the received video data. For example, image processing software algorithms such as OpenCV are used for the analysis. The analysis method detects changes in the learner's facial expressions, movements, and line of sight, and evaluates the learner's level of concentration and fatigue based on these.

[0742] generation means

[0743] Server: Based on the analysis results and past learning data, a generative AI model (e.g., a generative AI model) is used to generate optimized feedback for the learner. For example, if a learner is having trouble solving an equation, an explanation or exercise tailored to the content is generated.

[0744] Providing means

[0745] Device: The generated feedback is sent to the learner's tablet device and displayed to them. Specific study advice and the next task to be tackled based on the analysis results are displayed on the tablet screen.

[0746] Specific examples

[0747] For example, if a user is learning how to solve mathematical equations, a camera captures the user's learning activities and the video data is sent to a server via a tablet device. The server analyzes the video data and detects that the user is making a calculation error in a specific step. It also identifies that the user has repeatedly made mistakes in this area based on past practice test results. Based on this, the server generates detailed explanations for the user on "typical mistakes in solving equations and how to correct them" and sends them to the tablet device. The user then checks the explanations on the tablet and works on practice problems to correct their mistakes.

[0748] Prompt Sentence Examples

[0749] "Analyze where a user makes mistakes in solving math equations and generate specific explanations and practice questions for review."

[0750] "Calculate the learner's concentration and fatigue levels from video data and provide appropriate learning advice."

[0751] The above is an embodiment of the present invention. This system enables learners to receive optimal learning support in real time, thereby improving the quality of their learning.

[0752] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0753] Step 1:

[0754] Filming learning activities

[0755] Camera: A camera installed on each student's desk constantly records their learning activities, capturing in high resolution the student's facial expressions, movements, and gaze in detail.

[0756] Input: Real-time video of the learner

[0757] Output: High-resolution video data of learning activities

[0758] Step 2:

[0759] Video data compression and transmission

[0760] Terminal: The tablet terminal compresses the video data received from the camera using the H.264 algorithm, and then transmits the compressed data to the server using a secure communication protocol (e.g., HTTPS).

[0761] Input: High-resolution video of learning activities

[0762] Output: Compressed video data

[0763] Step 3:

[0764] Video data analysis

[0765] Server: Analyzes the received compressed video data. Image processing algorithms such as OpenCV are used for the analysis. The server analyzes the learner's facial expressions, movements, and changes in gaze to evaluate the learner's level of concentration and fatigue. It also compares this with past learning data.

[0766] Input: Compressed video data, past learning data

[0767] Output: Evaluation results of learner's concentration, fatigue, strengths and weaknesses

[0768] Step 4:

[0769] Generate feedback

[0770] Server: Based on the analysis results, the server sends a prompt to the generative AI model. The prompt includes the learner's situation and the feedback they need. The generative AI model generates specific feedback (explanations, practice questions, advice, etc.) based on the prompt.

[0771] Input: Analysis results, prompt text

[0772] Output: Generated feedback

[0773] Step 5:

[0774] Providing Feedback

[0775] Terminal: The feedback sent from the server is displayed on the tablet terminal. The user checks the feedback displayed on the tablet. Examples include "typical mistakes in solving equations and how to correct them."

[0776] Input: Generated feedback

[0777] Output: Feedback display on tablet device

[0778] Examples:

[0779] As an example, consider the case where a user is learning how to solve a mathematical equation.

[0780] Step 1: The camera captures the user's learning activity and captures it as high-resolution video data.

[0781] Step 2: The terminal compresses the video data using the H.264 algorithm and sends it to the server using the HTTPS protocol.

[0782] Step 3: The server analyzes the received video data using OpenCV to detect if the user is making an incorrect calculation at a specific step. It also integrates past learning data and evaluates the learner's strengths and weaknesses based on this.

[0783] Step 4: Based on the analysis results, the generative AI model is prompted to "analyze the user's mistakes in solving mathematical equations and generate specific explanations and practice questions for review." The generative AI model then generates specific feedback.

[0784] Step 5: The generated feedback is sent to the device, and "typical mistakes in solving equations and how to correct them" are displayed on the user's tablet. The user then tries learning again while looking at the results.

[0785] The above are the specific processing steps of the system of the present invention, which enable learners to receive optimal learning support in real time.

[0786] (Application example 1)

[0787] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0788] Conventional automated equipment operating in factories lacks real-time analysis functions to optimize efficiency and precision, and is unable to provide immediate feedback on improvements. As a result, waste and malfunctions tend to occur in the equipment's operation, leading to problems such as reduced productivity and poor quality.

[0789] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0790] In this invention, the server includes a photographing means for photographing the learning activities of the learner, a transmitting means for transmitting video data photographed by the photographing means to the server, an analyzing means for analyzing the received video data, a generating means for generating feedback appropriate for the learner based on the analysis results, a providing means for providing the generated feedback to the learner, an automation device operation analyzing means for analyzing the operation of automation devices operating in the factory and providing feedback on areas for improvement in efficiency and accuracy, and a display means for displaying the feedback. This makes it possible to monitor the operation of automation devices operating in the factory in real time and improve efficiency and accuracy based on the analysis results.

[0791] "Photographing means" refers to a camera or other image capturing device that captures the learning activities of learners and the operation of automated equipment in a factory.

[0792] The "transmission means" refers to a device or technology for transmitting the video data captured by the image capture means to the server, and is, for example, a device that uses a tablet terminal or a communication protocol.

[0793] The "analysis means" is a technology for analyzing the video data received in the server and evaluating the activities of the learners and the operation status of the automated devices.

[0794] "Generation means" refers to technology for generating appropriate feedback to learners and improvements to automated devices based on the analysis results.

[0795] "Providing means" refers to the devices and technologies used to display the generated feedback to the learner or operator.

[0796] "Automated equipment operation analysis means" is a technology for analyzing the operation of automated equipment operating in a factory and detecting areas for improvement in efficiency and accuracy.

[0797] "Display means" refers to a device or technology for visually presenting the analysis and generated feedback to the user, such as a tablet device or display.

[0798] "Real-time" refers to a time frame in which data is acquired, transmitted, analyzed, and feedback generated nearly simultaneously.

[0799] "Efficiency" is an index that indicates the degree to which automated equipment operating in a factory operates optimally and without waste.

[0800] "Accuracy" is a measure of the ability or degree to which an automated device accurately performs a required operation.

[0801] The present invention is a system for optimizing automated equipment in a factory by analyzing its operation and providing feedback on areas for improvement in efficiency and accuracy. This system is realized through the following steps.

[0802] System Program

[0803] This system mainly comprises an imaging means, a transmission means, an analysis means, a generation means, and a provision means.

[0804] Filming method

[0805] A camera is used as a recording medium. The camera records the operation of the automated equipment in real time and acquires the video data. The camera is connected to a tablet device via Wi-Fi or a wired connection.

[0806] Transmission method

[0807] The device receives the captured video data, compresses it, and transmits it to a server using a secure communication protocol. The high-resolution video data can capture even the smallest details of the automated equipment's operation.

[0808] Analysis means

[0809] The server analyzes the received video data using specific software algorithms (e.g., computer vision technology) to evaluate the efficiency and accuracy of the automated equipment's operation. This analysis identifies waste and malfunctions in the equipment's operation.

[0810] generation means

[0811] Based on the analysis results, feedback is generated to improve efficiency and accuracy. A generative AI model is used in this process to generate specific improvements and supplemental materials. For example, feedback such as "The movement of the right arm is slow and needs adjustment" is generated.

[0812] Providing means

[0813] The generated feedback is sent to the tablet device and displayed to the user. The feedback is provided in a visual format and includes specific instructions on how to improve.

[0814] Hardware and software used

[0815] Camera: Films the operation of automated equipment.

[0816] Tablet device: Receives video data captured by the camera and sends it to the server.

[0817] Server: Analyzes video data and generates feedback using computer vision techniques and generative AI models.

[0818] Specific examples

[0819] For example, if there is an automated machine that assembles a product, a camera will record its operation. The server will analyze the video data and identify the problem: "The right arm is moving slowly." Based on the analysis results, the generative AI model will generate feedback such as "Please adjust the right arm to move 10% faster," and send this to a tablet device. The user will then check this feedback on the tablet and adjust the machine's operation.

[0820] Prompt Sentence Examples

[0821] Configure your robot to analyze footage of its movements and identify areas for improvement in efficiency and accuracy. Analyze the movements in the footage and generate feedback that is user-friendly and specific.

[0822] This completes the detailed description of the embodiment of the present invention. This system makes it possible to monitor the operation of automated devices operating in a factory in real time and improve efficiency and accuracy based on the analysis results.

[0823] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0824] Step 1:

[0825] Camera image capture

[0826] Input: Factory automation equipment operation

[0827] Output: High-resolution video data

[0828] The camera captures the operation of the automated equipment in real time, capturing the operating status of the automated equipment and sending the video to a terminal as high-resolution video data.

[0829] Step 2:

[0830] Receive video data on the device

[0831] Input: High-resolution video data from a camera

[0832] Output: Compressed video data stored on the device

[0833] The device receives the video data captured by the camera and compresses it, making it ready to be sent to the server.

[0834] Step 3:

[0835] Compressed video data sent to server

[0836] Input: Compressed video data

[0837] Output: Video data sent to the server

[0838] The device transmits the compressed video data to the server using a secure communication protocol, with encryption techniques used to ensure data integrity and privacy.

[0839] Step 4:

[0840] Video data analysis by server

[0841] Input: Received video data

[0842] Output: Evaluation of the efficiency and accuracy of the operation as a result of the analysis

[0843] The server analyzes the received video data using computer vision techniques and performs calculations on the data to evaluate the operational status of the automated device (e.g., speed, accuracy of operation).

[0844] Step 5:

[0845] Feedback Generation

[0846] Input: Analysis results

[0847] Output: Feedback on areas for improvement in efficiency and accuracy

[0848] The server generates feedback based on the analysis results to improve efficiency and accuracy, using a generative AI model to generate specific improvements and supplementary materials.

[0849] Step 6:

[0850] Provide feedback to the device

[0851] Input: Generated feedback

[0852] Output: Feedback displayed on the terminal

[0853] The server sends the generated feedback to the device, which receives it and visually displays it to the user, allowing the user to confirm and implement specific improvements.

[0854] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0855] The present invention is a system for supporting learners in their learning, and is characterized by its ability to provide feedback based on the learner's emotional state by incorporating an emotion engine. An embodiment of this system will be described in detail.

[0856] This system mainly consists of the following components: a photographing means, a transmitting means, an analyzing means, a generating means, a providing means, and an emotion engine.

[0857] Filming method

[0858] Camera: A camera installed on each student's desk constantly records their learning activities. The camera is connected to a tablet device via Wi-Fi and captures video data in real time.

[0859] Transmission method

[0860] Terminal: The tablet terminal receives the video data captured by the camera, compresses it, and sends it to the server using a secure communication protocol.

[0861] Analysis means

[0862] Server: The server analyzes the received video data. Specific software algorithms are used to analyze the learner's facial expressions, movements, and changes in gaze. This analysis evaluates the learner's level of concentration and fatigue. It also integrates and analyzes past learning data (class notes, past questions, mock test results, etc.) to identify the learner's strengths and weaknesses.

[0863] Emotion Engine

[0864] Server: The emotion engine analyzes video and audio data to recognize the learner's emotional state. Specifically, it identifies the learner's emotions, such as joy, sadness, and stress, from facial expression and audio data. This emotional data is integrated with data obtained by the analysis means, enabling more accurate learning support.

[0865] generation means

[0866] Server: Generates optimized feedback for learners based on the analysis results and emotional data obtained by the emotion engine. A generative AI model is used in this process to generate specific advice, supplementary materials, additional questions, etc. For example, if a learner is feeling stressed, encouraging words and advice on how to relax are provided.

[0867] Providing means

[0868] Device: The generated feedback is sent to a tablet device and displayed to the user. Specific study advice and next tasks based on the analysis results and emotional state are displayed on the tablet screen. The user can confirm this and continue studying.

[0869] Specific examples

[0870] For example, consider a user learning how to solve a mathematical equation. A camera captures the user's learning, and the video data is sent to a server via a tablet device. The server analyzes the video data and detects that the user is making a calculation error at a specific step. Furthermore, the emotion engine recognizes from the user's facial expressions and voice that the user is feeling frustrated because they cannot solve the equation.

[0871] Based on this, the server generates detailed explanations for the user about "common mistakes in solving equations and how to correct them." It also adds encouraging words to the explanations, such as "This part is difficult, but try to stay calm. It's also recommended to take breaks." The generated feedback is sent to a tablet device, where the user can review it and work on practice problems to correct their mistakes. This allows the user to receive specific, emotionally appropriate feedback in real time.

[0872] As described above, the system of the present invention captures and analyzes learners' learning activities in real time and provides feedback that takes into account their emotional state, thereby realizing high-quality, individually optimized learning support.

[0873] The processing flow will be explained below.

[0874] Step 1:

[0875] User: Starts the learning application on the tablet device and configures the camera connection settings.

[0876] Step 2:

[0877] Device: The tablet device connects to the tabletop camera via Wi-Fi and checks the camera's connection status.

[0878] Step 3:

[0879] Camera: The tabletop camera continuously captures the learner's learning activities, capturing images at a set frame rate and resolution.

[0880] Step 4:

[0881] Terminal: Compresses video data captured in real time and sends it to the server using a secure communication protocol.

[0882] Step 5:

[0883] Server: Receives compressed video data and temporarily stores it in a database.

[0884] Step 6:

[0885] Server: Analyzes the received video data using the "analysis means." Specifically, it analyzes the learner's facial expressions, movements, and changes in line of sight to evaluate the learner's level of concentration and fatigue.

[0886] Step 7:

[0887] Server: Obtains past learning data (class notes, past questions, mock test results, etc.) from a database and integrates it with current analysis results to identify the learner's strengths and weaknesses.

[0888] Step 8:

[0889] Server: Using an "emotion engine," it analyzes video and audio data to recognize the learner's emotional state (happiness, sadness, stress, etc.). Specifically, it runs an algorithm that identifies emotions from facial expression and audio data.

[0890] Step 9:

[0891] Server: The "generator" generates optimal feedback based on the analysis results and emotional data. For example, if a calculation error is detected, it will include detailed instructions on how to correct the error and advice corresponding to the user's emotional state.

[0892] Step 10:

[0893] Server: Sends the generated feedback to the tablet device.

[0894] Step 11:

[0895] Device: The tablet device receives the feedback and displays it in the application's user interface.

[0896] Step 12:

[0897] User: Check the feedback displayed on the tablet and learn what to study next or what explanations to give. If necessary, follow the advice to relax.

[0898] Step 13:

[0899] Server: Records new progress and emotion data of the learner and stores it in a database. Based on the progress and emotion data, the learning plan and feedback content are updated.

[0900] Step 14:

[0901] User: Check their learning progress and emotional state, continue learning according to the plan, and repeat the cycle from step 1 if necessary.

[0902] Example 2

[0903] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0904] Conventional learning support systems can record learners' learning activities and provide feedback, but they have limitations in providing real-time support that takes into account the learner's emotional state. This makes it difficult to provide feedback appropriate to a learner's emotional state when they feel stressed or fatigued. Furthermore, there is a need for systems that go beyond simple data analysis to automatically identify a learner's strengths and weaknesses and provide individually optimized learning support based on that information.

[0905] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0906] In this invention, the server includes a photographing means for photographing the learning activities of the learner, a transmitting means for transmitting the video data photographed by the photographing means to the server, an analyzing means in the server for analyzing the received video data, a generating means for generating appropriate feedback based on the data and emotional data obtained by the analyzing means, a providing means for providing the generated feedback to the learner, an emotion recognition means for analyzing the emotional state of the learner from the video data and audio data, and a generating means using a generative AI model for integrating the obtained emotional data and data from the analyzing means to generate feedback optimized for the learner. This enables high-quality learning support that is individually optimized in real time and takes into account the emotional state of the learner.

[0907] "Photography means" refers to devices used to record learners' learning activities, such as cameras.

[0908] The "transmission means" refers to a device or function for transmitting the video data captured by the image capture means to the server.

[0909] "Analysis means" refers to software algorithms or programs used to analyze received video data and analyse changes in the learner's facial expressions, movements and gaze.

[0910] The "generation means" refers to a process or device that generates optimized feedback for the learner based on the data and emotional data obtained by the analysis means.

[0911] "Providing means" refers to a device or method for presenting the generated feedback to the learner.

[0912] "Emotion recognition means" refers to software or algorithms that analyze video and audio data to identify a learner's emotional state.

[0913] "Generative AI model" refers to an artificial intelligence model used to generate optimal feedback and advice for learners based on collected data.

[0914] "Feedback" refers to specific advice, supplementary materials, additional questions, words of encouragement, etc., generated by the analysis means or generation means and provided to the learner.

[0915] "Learning progress management means" refers to a process or device for recording a learner's progress and adjusting the learning plan based on the saved progress data.

[0916] The present invention is a system that comprehensively supports learners' learning activities, and its distinctive feature is that it incorporates the learner's emotional state. A specific embodiment of this system will be described. The main components include an image capture means, a transmission means, an analysis means, a generation means, a provision means, and an emotion recognition means. This makes it possible to provide learners with individually optimized feedback in real time.

[0917] First, when a user begins studying, a camera installed on the desk automatically starts operating. The camera captures the user's study activities in real time and continuously acquires video data. The camera is connected to a tablet device via Wi-Fi, and the captured video data is sent to the tablet device.

[0918] A camera connected to a terminal sends captured video data to the terminal. The terminal compresses the received video data and sends it to a server via a secure communication protocol (e.g., HTTPS). Data compression minimizes communication delays.

[0919] The server analyzes the received video data. Specifically, it uses OpenCV and other tools to analyze the user's facial expressions, movements, and changes in gaze. This allows it to evaluate the user's level of concentration and fatigue. It also integrates and analyzes past learning data (class notes, past questions, mock test results, etc.) to identify the user's strengths and weaknesses.

[0920] The emotion recognition means analyzes video and audio data to recognize the user's emotional state. Specifically, it identifies emotions from audio data using the Google Cloud Speech-to-Text API and integrates them with facial expression data. The results of this analysis are used to provide more accurate learning support.

[0921] The server generates optimized feedback for the user based on the acquired emotional data and learning data. Specifically, it uses a generative AI model (e.g., GPT-4) to generate feedback such as advice, supplementary materials, additional problems, and words of encouragement. For example, if a user feels stressed while working on a math problem, it generates a message recommending "how to approach the problem calmly" or "taking a break."

[0922] An example of a prompt to input to a generative AI model is as follows:

[0923] "You recognize that the user is getting fatigued while solving a math problem. Generate feedback that includes the best relaxation techniques for him / her and some encouraging words."

[0924] The device receives the generated feedback from the server and displays it on the user's tablet screen. Specific study advice and the next task to tackle based on the analysis results and emotional state are displayed on the screen. The user can review this and proceed with their studies based on it. For example, a message might be displayed saying, "This part is difficult, but try to stay calm. It's also recommended that you take a break."

[0925] As described above, the system of the present invention captures and analyzes learners' learning activities in real time and provides feedback that takes into account their emotional state, thereby realizing high-quality, individually optimized learning support.

[0926] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0927] Step 1:

[0928] When a user begins learning, a camera installed on the desk automatically starts operating. The camera captures the user's learning activities in real time and constantly acquires video data. The input is the user's learning activities, and the output is video data. This data is used in the next step.

[0929] Step 2:

[0930] A camera connected to a terminal receives captured video data. The received video data is compressed and sent to a server using a secure communication protocol (e.g., HTTPS). The input is video data from the camera, and the output is compressed video data sent to the server. This process minimizes data communication delays.

[0931] Step 3:

[0932] The server analyzes the received video data using a specific software algorithm (e.g., OpenCV). The input here is the received video data, and the output is an analysis result showing changes in the user's facial expressions, movements, and gaze. Specifically, facial recognition and gaze tracking within the video data are performed. Based on this analysis, the user's concentration and fatigue levels are evaluated.

[0933] Step 4:

[0934] The server analyzes the user's emotional state using emotion recognition. It receives video and audio data as input, and obtains emotional data as output. Specifically, it identifies emotions from audio data using the Google Cloud Speech-to-Text API and integrates this with facial expression data. For example, it can read stress or excitement from the tone and tempo of the user's voice.

[0935] Step 5:

[0936] The server generates feedback based on the acquired emotion data and analysis data. It takes the analysis results and emotion data as input, and generates feedback as output using a generative AI model (e.g., GPT-4). An example of a prompt sentence to input to the generative AI model is as follows:

[0937] "You recognize that the user is getting fatigued while solving a math problem. Generate feedback that includes the best relaxation techniques for him / her and some encouraging words."

[0938] This step generates feedback that may include specific advice, supplementary materials, additional questions, or words of encouragement.

[0939] Step 6:

[0940] The device receives the generated feedback sent from the server and displays it on the user's tablet screen. The input is feedback data from the server, and the output is specific study advice and the next task the user should tackle. The user can then proceed with their studies based on this. For example, a message such as "This part is difficult, but try to stay calm. It is also recommended that you take a break" may be displayed.

[0941] (Application example 2)

[0942] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0943] Conventional learning support systems focus on assessing learners' levels of concentration and fatigue, but they also need to provide appropriate feedback in real time, not only to learners but also in other environments (e.g., customer service in brick-and-mortar stores). Such systems require advanced analytical techniques that take into account human emotions and behavior, which has been difficult to achieve with conventional technologies. It has also been difficult to develop a system that can accurately grasp a customer's emotional state, such as interest or confusion, in a store and provide appropriate advice on the spot.

[0944] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for capturing and evaluating the behavior and facial expressions of customers in the store, means including a generative AI model for generating optimal advice for the customer based on the evaluation results, and means for evaluating the customer's interest and confusion and generating appropriate prompt sentences. This makes it possible to analyze the customer's emotions and behavior in real time and provide appropriate advice on the spot.

[0945] A "learner" is an individual who uses a learning support system to carry out learning activities.

[0946] A "customer" is an individual who expresses interest in a product or service within a physical store.

[0947] "Photography means" refers to cameras or smart glasses used to capture the behavior and expressions of learners and customers.

[0948] The "transmission means" is a communication device for transmitting the video data captured by the image capture means to the server.

[0949] The "analysis means" refers to software and algorithms that analyze the video data received by the server and evaluate the behavior and emotions of learners and customers.

[0950] "Generation means" refers to a generative AI model that generates feedback appropriate for learners and customers based on the analysis results.

[0951] The "delivery means" is a display or notification system for providing the generated feedback to learners or customers.

[0952] A "generative AI model" is an artificial intelligence model that automatically generates optimal feedback and advice based on analysis results.

[0953] A "prompt" is a sentence that is input into a generative AI model and is an instruction sentence used to generate feedback or advice.

[0954] This invention is a system that can be applied not only to learners' learning activities but also to customer service support in brick-and-mortar stores. The system of the present invention is characterized by analyzing the behavior, facial expressions, and emotional states of learners and customers in real time and providing appropriate feedback and advice.

[0955] Hardware and Software

[0956] The main hardware used in this system is as follows:

[0957] Capture method: camera or smart glasses (e.g. Google Glass)

[0958] Transmission method: Tablet device or smartphone

[0959] Server: High-performance server (e.g. AWS EC2 instance)

[0960] The software used includes:

[0961] Video data analysis software: OpenCV

[0962] Sentiment analysis engine: Microsoft Azure Emotion API

[0963] Generative AI model: OpenAI's GPT-3 model

[0964] Data processing and calculation

[0965] The server performs the following steps:

[0966] 1. Acquisition of video data: Video data acquired by a camera is sent to a terminal. This data includes the behavior and facial expressions of the learner or customer.

[0967] 2. Data transmission: The device transmits the captured video data to the server, using data compression and a secure communication protocol.

[0968] 3. Video analysis and emotion assessment: The server analyzes the video data using OpenCV and assesses the emotional state using the Microsoft Azure Emotion API. From the analysis results, it identifies the learner's concentration level or fatigue level, and the customer's interest or confusion.

[0969] 4. Feedback generation: Based on the analysis results, the generative AI model (GPT-3) generates specific feedback and advice, including study advice and product information.

[0970] 5. Providing advice: The generated feedback is displayed and notified on the tablet device or smart glasses that are used as the means of providing the advice.

[0971] Examples and prompts

[0972] For example, in a brick-and-mortar scenario, consider a situation where a customer is interested in a particular product but is confused by its features. The server can analyze this confusion from the customer's facial expressions and behavior, and then use a generative AI model to input the following prompt:

[0973] "A customer was observed on camera. They were looking at product X with interest, but seemed confused about its features. Please generate additional information to provide to the customer."

[0974] Here, the generative AI model generates specific product descriptions and advice, which are displayed on the smart glasses, allowing store staff to provide information to customers at the right time.

[0975] In this way, a system can be built that provides optimal feedback to customers and learners in real time, achieving high-quality service and learning support.

[0976] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0977] Step 1:

[0978] The user begins learning activities or in-store activities. The imaging means (camera or smart glasses) captures the user's actions and facial expressions in real time. The input is the user's learning activities or in-store activities and facial expressions, and the output is video data.

[0979] Step 2:

[0980] The captured video data is sent to a device (tablet or smartphone). The device compresses the video data and sends it to a server using a secure communication protocol. The input is uncompressed video data, and the output is compressed video data.

[0981] Step 3:

[0982] The server analyzes the received video data. It processes visual information using OpenCV and evaluates emotional states using the Microsoft Azure Emotion API. From the analysis results, it identifies the user's behavior, facial expressions, concentration level, fatigue level, and customer interest or confusion. The input is compressed video data, and the output is analyzed behavioral and emotional data.

[0983] Step 4:

[0984] Based on the analysis results, the server uses a generative AI model (GPT-3) to generate specific feedback and advice. The prompt sentence is input into the generative AI model to obtain feedback appropriate for the specific situation. The input is the analyzed behavioral data, emotional data, and the prompt sentence, and the output is specific feedback and advice.

[0985] Step 5:

[0986] The means of providing this feedback (tablet device or smart glasses) displays the generated feedback and advice to the user, who then modifies or continues their learning or behavior accordingly. The input is the generated feedback and advice, and the output is the feedback provided to the user.

[0987] This process allows users to receive appropriate feedback in real time, enabling efficient learning and a meaningful purchasing experience.

[0988] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0989] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0990] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[0991] [Fourth embodiment]

[0992] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[0993] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0994] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0995] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[0996] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0997] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0998] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0999] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1000] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1001] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1002] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1003] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1004] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1005] The present invention is a system for supporting learners' learning, specifically, a system that analyzes and supports learners' activities in real time using cameras and tablet devices installed in the learning environment. An embodiment of this system will be described in detail.

[1006] This system mainly comprises the following components: an imaging means, a transmitting means, an analyzing means, a generating means, and a providing means.

[1007] Filming method

[1008] Camera: A camera installed on each student's desk constantly records their learning activities. The camera is connected to a tablet device via Wi-Fi and captures video data in real time.

[1009] Transmission method

[1010] Device: The tablet device receives the video data captured by the camera, compresses it, and sends it to the server using a secure communication protocol. The video data is high resolution and can capture the learner's detailed movements and facial expressions.

[1011] Analysis means

[1012] Server: The server analyzes the received video data. Specific software algorithms are used to analyze the learner's facial expressions, movements, and changes in gaze. This analysis makes it possible to assess the learner's level of concentration and fatigue. Furthermore, by integrating and analyzing past learning data (class notes, past questions, mock test results, etc.), the server can identify the learner's strengths and weaknesses.

[1013] generation means

[1014] Server: Based on the analysis results, it generates optimized feedback for the learner. A generative AI model is used in this process to generate specific advice, supplementary materials, additional problems, etc. For example, if a learner is having trouble solving an equation, an explanation or practice problem will be provided to encourage review of that part.

[1015] Providing means

[1016] Device: The generated feedback is sent to a tablet device and displayed to the user. Specific study advice and the next task to tackle based on the analysis results are displayed on the tablet screen. The user can check this and continue studying.

[1017] Specific examples

[1018] For example, consider a user learning how to solve mathematical equations. A camera captures the user's learning, and the video data is sent to a server via a tablet device. The server analyzes the video data and detects that the user is making a calculation error in a particular step. It also identifies that the user has repeatedly made mistakes in this area based on past practice test results.

[1019] Based on this, the server generates detailed explanations for the user on "typical mistakes in solving equations and how to correct them" and sends them to the tablet device. The user can then review the explanations on the tablet and work on practice problems to correct the mistakes. This allows the user to receive specific, individually optimized learning support in real time.

[1020] As described above, the system of the present invention realizes high-quality learning support by capturing and analyzing learners' learning activities in real time and providing individually optimized feedback.

[1021] The processing flow will be explained below.

[1022] Step 1:

[1023] User: Starts the learning application on the tablet device and configures the camera connection settings.

[1024] Step 2:

[1025] Device: The tablet device connects to the tabletop camera via Wi-Fi and checks the camera's connection status.

[1026] Step 3:

[1027] Camera: The tabletop camera continuously captures the learner's learning activities, capturing images at a set frame rate and resolution.

[1028] Step 4:

[1029] Terminal: Compresses video data captured in real time and sends it to the server using a secure communication protocol.

[1030] Step 5:

[1031] Server: Receives compressed video data and temporarily stores it in a database.

[1032] Step 6:

[1033] Server: Using "Gemini," the received video data is analyzed. Specifically, changes in the learner's facial expressions, movements, and gaze are analyzed to evaluate the learner's level of concentration and fatigue.

[1034] Step 7:

[1035] Server: Based on the analysis results, the learner's strengths and weaknesses are identified, and past learning data (class notes, past questions, mock test results, etc.) is integrated as necessary for analysis.

[1036] Step 8:

[1037] Server: Based on the analysis results, a generative AI model generates optimized feedback for the learner, including specific advice, supplementary materials, and additional questions.

[1038] Step 9:

[1039] Server: Sends the generated feedback to the tablet device.

[1040] Step 10:

[1041] Device: The tablet device receives the feedback and displays it in the application's user interface.

[1042] Step 11:

[1043] User: Check the feedback displayed on the tablet and learn what to study next or what explanations to give.

[1044] Step 12:

[1045] Server: Records new progress data of learners and stores it in a database. Based on the progress, the learning plan and feedback are updated.

[1046] Step 13:

[1047] User: Check the learning progress and continue learning according to the plan. If necessary, repeat the cycle from step 1.

[1048] Example 1

[1049] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1050] Conventional learning support systems have difficulty analyzing learners' learning activities in detail in real time, making it difficult to provide individually optimized feedback in a timely manner. Furthermore, they lack the functionality to capture subtle changes in learners' facial expressions and gaze to evaluate their concentration and fatigue levels, making it difficult to optimally improve learners' performance.

[1051] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1052] In this invention, the server includes a photographing means, a transmitting means, an analyzing means, a generating means, and a providing means. This enables the learner's learning activities to be analyzed in real time and individually optimized feedback to be provided promptly. Specifically, the server uses a specific software algorithm to analyze the received video data and analyzes the learner's facial expressions, movements, and changes in gaze to evaluate the learner's level of concentration and fatigue. Based on the results, the server uses a generative AI model to generate feedback appropriate for the learner, and has a means for transmitting and displaying the feedback on the mobile information terminal. This allows the learner to receive optimal learning support in real time, improving the quality of their learning.

[1053] The "filming means" is a device that films the learning activities of the learner in real time and acquires video data.

[1054] The "transmission means" is a device or software that has the function of compressing the video data acquired by the image capture means and transmitting it to the server using a secure communication protocol.

[1055] "Analysis means" refers to a device or software that has the function of analyzing the video data received by the server using a specific software algorithm, and analyzing changes in the learner's facial expressions, movements, and gaze to evaluate their level of concentration and fatigue.

[1056] "Generation means" refers to a device or software that has the function of utilizing a generative AI model to automatically generate feedback optimized for the learner based on the analysis results obtained by the analysis means and past learning data.

[1057] The "providing means" is a device or software that has the function of transmitting the feedback generated by the generating means to the learner's mobile information terminal and displaying it.

[1058] A "generative AI model" is a model that uses artificial intelligence algorithms to generate appropriate feedback based on input data.

[1059] This system is constructed to analyze learners' learning activities in detail in real time and provide individually optimized feedback. The system consists of a capturing means, a transmitting means, an analyzing means, a generating means, and a providing means.

[1060] Filming method

[1061] Camera: A camera (e.g., a high-resolution webcam) installed on the student's desk continuously records the student's learning activities. The camera is connected to a tablet device via Wi-Fi and captures video data in real time.

[1062] Transmission method

[1063] Terminal: The tablet terminal receives the video data captured by the camera and automatically compresses it using an efficient video compression algorithm such as H.264. The compressed video data is then sent to the server using a secure communication protocol (e.g., HTTPS).

[1064] Analysis means

[1065] Server: The server analyzes the received video data. For example, image processing software algorithms such as OpenCV are used for the analysis. The analysis method detects changes in the learner's facial expressions, movements, and line of sight, and evaluates the learner's level of concentration and fatigue based on these.

[1066] generation means

[1067] Server: Based on the analysis results and past learning data, a generative AI model (e.g., a generative AI model) is used to generate optimized feedback for the learner. For example, if a learner is having trouble solving an equation, an explanation or exercise tailored to the content is generated.

[1068] Providing means

[1069] Device: The generated feedback is sent to the learner's tablet device and displayed to them. Specific study advice and the next task to be tackled based on the analysis results are displayed on the tablet screen.

[1070] Specific examples

[1071] For example, if a user is learning how to solve mathematical equations, a camera captures the user's learning activities and the video data is sent to a server via a tablet device. The server analyzes the video data and detects that the user is making a calculation error in a specific step. It also identifies that the user has repeatedly made mistakes in this area based on past practice test results. Based on this, the server generates detailed explanations for the user on "typical mistakes in solving equations and how to correct them" and sends them to the tablet device. The user then checks the explanations on the tablet and works on practice problems to correct their mistakes.

[1072] Prompt Sentence Examples

[1073] "Analyze where a user makes mistakes in solving math equations and generate specific explanations and practice questions for review."

[1074] "Calculate the learner's concentration and fatigue levels from video data and provide appropriate learning advice."

[1075] The above is an embodiment of the present invention. This system enables learners to receive optimal learning support in real time, thereby improving the quality of their learning.

[1076] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1077] Step 1:

[1078] Filming learning activities

[1079] Camera: A camera installed on each student's desk constantly records their learning activities, capturing in high resolution the student's facial expressions, movements, and gaze in detail.

[1080] Input: Real-time video of the learner

[1081] Output: High-resolution video data of learning activities

[1082] Step 2:

[1083] Video data compression and transmission

[1084] Terminal: The tablet terminal compresses the video data received from the camera using the H.264 algorithm, and then transmits the compressed data to the server using a secure communication protocol (e.g., HTTPS).

[1085] Input: High-resolution video of learning activities

[1086] Output: Compressed video data

[1087] Step 3:

[1088] Video data analysis

[1089] Server: Analyzes the received compressed video data. Image processing algorithms such as OpenCV are used for the analysis. The server analyzes the learner's facial expressions, movements, and changes in gaze to evaluate the learner's level of concentration and fatigue. It also compares this with past learning data.

[1090] Input: Compressed video data, past learning data

[1091] Output: Evaluation results of learner's concentration, fatigue, strengths and weaknesses

[1092] Step 4:

[1093] Generate feedback

[1094] Server: Based on the analysis results, the server sends a prompt to the generative AI model. The prompt includes the learner's situation and the feedback they need. The generative AI model generates specific feedback (explanations, practice questions, advice, etc.) based on the prompt.

[1095] Input: Analysis results, prompt text

[1096] Output: Generated feedback

[1097] Step 5:

[1098] Providing Feedback

[1099] Terminal: The feedback sent from the server is displayed on the tablet terminal. The user checks the feedback displayed on the tablet. Examples include "typical mistakes in solving equations and how to correct them."

[1100] Input: Generated feedback

[1101] Output: Feedback display on tablet device

[1102] Examples:

[1103] As an example, consider the case where a user is learning how to solve a mathematical equation.

[1104] Step 1: The camera captures the user's learning activity and captures it as high-resolution video data.

[1105] Step 2: The terminal compresses the video data using the H.264 algorithm and sends it to the server using the HTTPS protocol.

[1106] Step 3: The server analyzes the received video data using OpenCV to detect if the user is making an incorrect calculation at a specific step. It also integrates past learning data and evaluates the learner's strengths and weaknesses based on this.

[1107] Step 4: Based on the analysis results, the generative AI model is prompted to "analyze the user's mistakes in solving mathematical equations and generate specific explanations and practice questions for review." The generative AI model then generates specific feedback.

[1108] Step 5: The generated feedback is sent to the device, and "typical mistakes in solving equations and how to correct them" are displayed on the user's tablet. The user then tries learning again while looking at the results.

[1109] The above are the specific processing steps of the system of the present invention, which enable learners to receive optimal learning support in real time.

[1110] (Application example 1)

[1111] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1112] Conventional automated equipment operating in factories lacks real-time analysis functions to optimize efficiency and precision, and is unable to provide immediate feedback on improvements. As a result, waste and malfunctions tend to occur in the equipment's operation, leading to problems such as reduced productivity and poor quality.

[1113] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1114] In this invention, the server includes a photographing means for photographing the learning activities of the learner, a transmitting means for transmitting video data photographed by the photographing means to the server, an analyzing means for analyzing the received video data, a generating means for generating feedback appropriate for the learner based on the analysis results, a providing means for providing the generated feedback to the learner, an automation device operation analyzing means for analyzing the operation of automation devices operating in the factory and providing feedback on areas for improvement in efficiency and accuracy, and a display means for displaying the feedback. This makes it possible to monitor the operation of automation devices operating in the factory in real time and improve efficiency and accuracy based on the analysis results.

[1115] "Photographing means" refers to a camera or other image capturing device that captures the learning activities of learners and the operation of automated equipment in a factory.

[1116] The "transmission means" refers to a device or technology for transmitting the video data captured by the image capture means to the server, and is, for example, a device that uses a tablet terminal or a communication protocol.

[1117] The "analysis means" is a technology for analyzing the video data received in the server and evaluating the activities of the learners and the operation status of the automated devices.

[1118] "Generation means" refers to technology for generating appropriate feedback to learners and improvements to automated devices based on the analysis results.

[1119] "Providing means" refers to the devices and technologies used to display the generated feedback to the learner or operator.

[1120] "Automated equipment operation analysis means" is a technology for analyzing the operation of automated equipment operating in a factory and detecting areas for improvement in efficiency and accuracy.

[1121] "Display means" refers to a device or technology for visually presenting the analysis and generated feedback to the user, such as a tablet device or display.

[1122] "Real-time" refers to a time frame in which data is acquired, transmitted, analyzed, and feedback generated nearly simultaneously.

[1123] "Efficiency" is an index that indicates the degree to which automated equipment operating in a factory operates optimally and without waste.

[1124] "Accuracy" is a measure of the ability or degree to which an automated device accurately performs a required operation.

[1125] The present invention is a system for optimizing automated equipment in a factory by analyzing its operation and providing feedback on areas for improvement in efficiency and accuracy. This system is realized through the following steps.

[1126] System Program

[1127] This system mainly comprises an imaging means, a transmission means, an analysis means, a generation means, and a provision means.

[1128] Filming method

[1129] A camera is used as a recording medium. The camera records the operation of the automated equipment in real time and acquires the video data. The camera is connected to a tablet device via Wi-Fi or a wired connection.

[1130] Transmission method

[1131] The device receives the captured video data, compresses it, and transmits it to a server using a secure communication protocol. The high-resolution video data can capture even the smallest details of the automated equipment's operation.

[1132] Analysis means

[1133] The server analyzes the received video data using specific software algorithms (e.g., computer vision technology) to evaluate the efficiency and accuracy of the automated equipment's operation. This analysis identifies waste and malfunctions in the equipment's operation.

[1134] generation means

[1135] Based on the analysis results, feedback is generated to improve efficiency and accuracy. A generative AI model is used in this process to generate specific improvements and supplemental materials. For example, feedback such as "The movement of the right arm is slow and needs adjustment" is generated.

[1136] Providing means

[1137] The generated feedback is sent to the tablet device and displayed to the user. The feedback is provided in a visual format and includes specific instructions on how to improve.

[1138] Hardware and software used

[1139] Camera: Films the operation of automated equipment.

[1140] Tablet device: Receives video data captured by the camera and sends it to the server.

[1141] Server: Analyzes video data and generates feedback using computer vision techniques and generative AI models.

[1142] Specific examples

[1143] For example, if there is an automated machine that assembles a product, a camera will record its operation. The server will analyze the video data and identify the problem: "The right arm is moving slowly." Based on the analysis results, the generative AI model will generate feedback such as "Please adjust the right arm to move 10% faster," and send this to a tablet device. The user will then check this feedback on the tablet and adjust the machine's operation.

[1144] Prompt Sentence Examples

[1145] Configure your robot to analyze footage of its movements and identify areas for improvement in efficiency and accuracy. Analyze the movements in the footage and generate feedback that is user-friendly and specific.

[1146] This completes the detailed description of the embodiment of the present invention. This system makes it possible to monitor the operation of automated devices operating in a factory in real time and improve efficiency and accuracy based on the analysis results.

[1147] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1148] Step 1:

[1149] Camera image capture

[1150] Input: Factory automation equipment operation

[1151] Output: High-resolution video data

[1152] The camera captures the operation of the automated equipment in real time, capturing the operating status of the automated equipment and sending the video to a terminal as high-resolution video data.

[1153] Step 2:

[1154] Receive video data on the device

[1155] Input: High-resolution video data from a camera

[1156] Output: Compressed video data stored on the device

[1157] The device receives the video data captured by the camera and compresses it, making it ready to be sent to the server.

[1158] Step 3:

[1159] Compressed video data sent to server

[1160] Input: Compressed video data

[1161] Output: Video data sent to the server

[1162] The device transmits the compressed video data to the server using a secure communication protocol, with encryption techniques used to ensure data integrity and privacy.

[1163] Step 4:

[1164] Video data analysis by server

[1165] Input: Received video data

[1166] Output: Evaluation of the efficiency and accuracy of the operation as a result of the analysis

[1167] The server analyzes the received video data using computer vision techniques and performs calculations on the data to evaluate the operational status of the automated device (e.g., speed, accuracy of operation).

[1168] Step 5:

[1169] Feedback Generation

[1170] Input: Analysis results

[1171] Output: Feedback on areas for improvement in efficiency and accuracy

[1172] The server generates feedback based on the analysis results to improve efficiency and accuracy, using a generative AI model to generate specific improvements and supplementary materials.

[1173] Step 6:

[1174] Provide feedback to the device

[1175] Input: Generated feedback

[1176] Output: Feedback displayed on the terminal

[1177] The server sends the generated feedback to the device, which receives it and visually displays it to the user, allowing the user to confirm and implement specific improvements.

[1178] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1179] The present invention is a system for supporting learners in their learning, and is characterized by its ability to provide feedback based on the learner's emotional state by incorporating an emotion engine. An embodiment of this system will be described in detail.

[1180] This system mainly consists of the following components: a photographing means, a transmitting means, an analyzing means, a generating means, a providing means, and an emotion engine.

[1181] Filming method

[1182] Camera: A camera installed on each student's desk constantly records their learning activities. The camera is connected to a tablet device via Wi-Fi and captures video data in real time.

[1183] Transmission method

[1184] Terminal: The tablet terminal receives the video data captured by the camera, compresses it, and sends it to the server using a secure communication protocol.

[1185] Analysis means

[1186] Server: The server analyzes the received video data. Specific software algorithms are used to analyze the learner's facial expressions, movements, and changes in gaze. This analysis evaluates the learner's level of concentration and fatigue. It also integrates and analyzes past learning data (class notes, past questions, mock test results, etc.) to identify the learner's strengths and weaknesses.

[1187] Emotion Engine

[1188] Server: The emotion engine analyzes video and audio data to recognize the learner's emotional state. Specifically, it identifies the learner's emotions, such as joy, sadness, and stress, from facial expression and audio data. This emotional data is integrated with data obtained by the analysis means, enabling more accurate learning support.

[1189] generation means

[1190] Server: Generates optimized feedback for learners based on the analysis results and emotional data obtained by the emotion engine. A generative AI model is used in this process to generate specific advice, supplementary materials, additional questions, etc. For example, if a learner is feeling stressed, encouraging words and advice on how to relax are provided.

[1191] Providing means

[1192] Device: The generated feedback is sent to a tablet device and displayed to the user. Specific study advice and next tasks based on the analysis results and emotional state are displayed on the tablet screen. The user can confirm this and continue studying.

[1193] Specific examples

[1194] For example, consider a user learning how to solve a mathematical equation. A camera captures the user's learning, and the video data is sent to a server via a tablet device. The server analyzes the video data and detects that the user is making a calculation error at a specific step. Furthermore, the emotion engine recognizes from the user's facial expressions and voice that the user is feeling frustrated because they cannot solve the equation.

[1195] Based on this, the server generates detailed explanations for the user about "common mistakes in solving equations and how to correct them." It also adds encouraging words to the explanations, such as "This part is difficult, but try to stay calm. It's also recommended to take breaks." The generated feedback is sent to a tablet device, where the user can review it and work on practice problems to correct their mistakes. This allows the user to receive specific, emotionally appropriate feedback in real time.

[1196] As described above, the system of the present invention captures and analyzes learners' learning activities in real time and provides feedback that takes into account their emotional state, thereby realizing high-quality, individually optimized learning support.

[1197] The processing flow will be explained below.

[1198] Step 1:

[1199] User: Starts the learning application on the tablet device and configures the camera connection settings.

[1200] Step 2:

[1201] Device: The tablet device connects to the tabletop camera via Wi-Fi and checks the camera's connection status.

[1202] Step 3:

[1203] Camera: The tabletop camera continuously captures the learner's learning activities, capturing images at a set frame rate and resolution.

[1204] Step 4:

[1205] Terminal: Compresses video data captured in real time and sends it to the server using a secure communication protocol.

[1206] Step 5:

[1207] Server: Receives compressed video data and temporarily stores it in a database.

[1208] Step 6:

[1209] Server: Analyzes the received video data using the "analysis means." Specifically, it analyzes the learner's facial expressions, movements, and changes in line of sight to evaluate the learner's level of concentration and fatigue.

[1210] Step 7:

[1211] Server: Obtains past learning data (class notes, past questions, mock test results, etc.) from a database and integrates it with current analysis results to identify the learner's strengths and weaknesses.

[1212] Step 8:

[1213] Server: Using an "emotion engine," it analyzes video and audio data to recognize the learner's emotional state (happiness, sadness, stress, etc.). Specifically, it runs an algorithm that identifies emotions from facial expression and audio data.

[1214] Step 9:

[1215] Server: The "generator" generates optimal feedback based on the analysis results and emotional data. For example, if a calculation error is detected, it will include detailed instructions on how to correct the error and advice corresponding to the user's emotional state.

[1216] Step 10:

[1217] Server: Sends the generated feedback to the tablet device.

[1218] Step 11:

[1219] Device: The tablet device receives the feedback and displays it in the application's user interface.

[1220] Step 12:

[1221] User: Check the feedback displayed on the tablet and learn what to study next or what explanations to give. If necessary, follow the advice to relax.

[1222] Step 13:

[1223] Server: Records new progress and emotion data of the learner and stores it in a database. Based on the progress and emotion data, the learning plan and feedback content are updated.

[1224] Step 14:

[1225] User: Check their learning progress and emotional state, continue learning according to the plan, and repeat the cycle from step 1 if necessary.

[1226] Example 2

[1227] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1228] Conventional learning support systems can record learners' learning activities and provide feedback, but they have limitations in providing real-time support that takes into account the learner's emotional state. This makes it difficult to provide feedback appropriate to a learner's emotional state when they feel stressed or fatigued. Furthermore, there is a need for systems that go beyond simple data analysis to automatically identify a learner's strengths and weaknesses and provide individually optimized learning support based on that information.

[1229] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1230] In this invention, the server includes a photographing means for photographing the learning activities of the learner, a transmitting means for transmitting the video data photographed by the photographing means to the server, an analyzing means in the server for analyzing the received video data, a generating means for generating appropriate feedback based on the data and emotional data obtained by the analyzing means, a providing means for providing the generated feedback to the learner, an emotion recognition means for analyzing the emotional state of the learner from the video data and audio data, and a generating means using a generative AI model for integrating the obtained emotional data and data from the analyzing means to generate feedback optimized for the learner. This enables high-quality learning support that is individually optimized in real time and takes into account the emotional state of the learner.

[1231] "Photography means" refers to devices used to record learners' learning activities, such as cameras.

[1232] The "transmission means" refers to a device or function for transmitting the video data captured by the image capture means to the server.

[1233] "Analysis means" refers to software algorithms or programs used to analyze received video data and analyse changes in the learner's facial expressions, movements and gaze.

[1234] The "generation means" refers to a process or device that generates optimized feedback for the learner based on the data and emotional data obtained by the analysis means.

[1235] "Providing means" refers to a device or method for presenting the generated feedback to the learner.

[1236] "Emotion recognition means" refers to software or algorithms that analyze video and audio data to identify a learner's emotional state.

[1237] "Generative AI model" refers to an artificial intelligence model used to generate optimal feedback and advice for learners based on collected data.

[1238] "Feedback" refers to specific advice, supplementary materials, additional questions, words of encouragement, etc., generated by the analysis means or generation means and provided to the learner.

[1239] "Learning progress management means" refers to a process or device for recording a learner's progress and adjusting the learning plan based on the saved progress data.

[1240] The present invention is a system that comprehensively supports learners' learning activities, and its distinctive feature is that it incorporates the learner's emotional state. A specific embodiment of this system will be described. The main components include an image capture means, a transmission means, an analysis means, a generation means, a provision means, and an emotion recognition means. This makes it possible to provide learners with individually optimized feedback in real time.

[1241] First, when a user begins studying, a camera installed on the desk automatically starts operating. The camera captures the user's study activities in real time and continuously acquires video data. The camera is connected to a tablet device via Wi-Fi, and the captured video data is sent to the tablet device.

[1242] A camera connected to a terminal sends captured video data to the terminal. The terminal compresses the received video data and sends it to a server via a secure communication protocol (e.g., HTTPS). Data compression minimizes communication delays.

[1243] The server analyzes the received video data. Specifically, it uses OpenCV and other tools to analyze the user's facial expressions, movements, and changes in gaze. This allows it to evaluate the user's level of concentration and fatigue. It also integrates and analyzes past learning data (class notes, past questions, mock test results, etc.) to identify the user's strengths and weaknesses.

[1244] The emotion recognition means analyzes video and audio data to recognize the user's emotional state. Specifically, it identifies emotions from audio data using the Google Cloud Speech-to-Text API and integrates them with facial expression data. The results of this analysis are used to provide more accurate learning support.

[1245] The server generates optimized feedback for the user based on the acquired emotional data and learning data. Specifically, it uses a generative AI model (e.g., GPT-4) to generate feedback such as advice, supplementary materials, additional problems, and words of encouragement. For example, if a user feels stressed while working on a math problem, it generates a message recommending "how to approach the problem calmly" or "taking a break."

[1246] An example of a prompt to input to a generative AI model is as follows:

[1247] "You recognize that the user is getting fatigued while solving a math problem. Generate feedback that includes the best relaxation techniques for him / her and some encouraging words."

[1248] The device receives the generated feedback from the server and displays it on the user's tablet screen. Specific study advice and the next task to tackle based on the analysis results and emotional state are displayed on the screen. The user can review this and proceed with their studies based on it. For example, a message might be displayed saying, "This part is difficult, but try to stay calm. It's also recommended that you take a break."

[1249] As described above, the system of the present invention captures and analyzes learners' learning activities in real time and provides feedback that takes into account their emotional state, thereby realizing high-quality, individually optimized learning support.

[1250] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1251] Step 1:

[1252] When a user begins learning, a camera installed on the desk automatically starts operating. The camera captures the user's learning activities in real time and constantly acquires video data. The input is the user's learning activities, and the output is video data. This data is used in the next step.

[1253] Step 2:

[1254] A camera connected to a terminal receives captured video data. The received video data is compressed and sent to a server using a secure communication protocol (e.g., HTTPS). The input is video data from the camera, and the output is compressed video data sent to the server. This process minimizes data communication delays.

[1255] Step 3:

[1256] The server analyzes the received video data using a specific software algorithm (e.g., OpenCV). The input here is the received video data, and the output is an analysis result showing changes in the user's facial expressions, movements, and gaze. Specifically, facial recognition and gaze tracking within the video data are performed. Based on this analysis, the user's concentration and fatigue levels are evaluated.

[1257] Step 4:

[1258] The server analyzes the user's emotional state using emotion recognition. It receives video and audio data as input, and obtains emotional data as output. Specifically, it identifies emotions from audio data using the Google Cloud Speech-to-Text API and integrates this with facial expression data. For example, it can read stress or excitement from the tone and tempo of the user's voice.

[1259] Step 5:

[1260] The server generates feedback based on the acquired emotion data and analysis data. It takes the analysis results and emotion data as input, and generates feedback as output using a generative AI model (e.g., GPT-4). An example of a prompt sentence to input to the generative AI model is as follows:

[1261] "You recognize that the user is getting fatigued while solving a math problem. Generate feedback that includes the best relaxation techniques for him / her and some encouraging words."

[1262] This step generates feedback that may include specific advice, supplementary materials, additional questions, or words of encouragement.

[1263] Step 6:

[1264] The device receives the generated feedback sent from the server and displays it on the user's tablet screen. The input is feedback data from the server, and the output is specific study advice and the next task the user should tackle. The user can then proceed with their studies based on this. For example, a message such as "This part is difficult, but try to stay calm. It is also recommended that you take a break" may be displayed.

[1265] (Application example 2)

[1266] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1267] Conventional learning support systems focus on assessing learners' levels of concentration and fatigue, but they also need to provide appropriate feedback in real time, not only to learners but also in other environments (e.g., customer service in brick-and-mortar stores). Such systems require advanced analytical techniques that take into account human emotions and behavior, which has been difficult to achieve with conventional technologies. It has also been difficult to develop a system that can accurately grasp a customer's emotional state, such as interest or confusion, in a store and provide appropriate advice on the spot.

[1268] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for capturing and evaluating the behavior and facial expressions of customers in the store, means including a generative AI model for generating optimal advice for the customer based on the evaluation results, and means for evaluating the customer's interest and confusion and generating appropriate prompt sentences. This makes it possible to analyze the customer's emotions and behavior in real time and provide appropriate advice on the spot.

[1269] A "learner" is an individual who uses a learning support system to carry out learning activities.

[1270] A "customer" is an individual who expresses interest in a product or service within a physical store.

[1271] "Photography means" refers to cameras or smart glasses used to capture the behavior and expressions of learners and customers.

[1272] The "transmission means" is a communication device for transmitting the video data captured by the image capture means to the server.

[1273] The "analysis means" refers to software and algorithms that analyze the video data received by the server and evaluate the behavior and emotions of learners and customers.

[1274] "Generation means" refers to a generative AI model that generates feedback appropriate for learners and customers based on the analysis results.

[1275] The "delivery means" is a display or notification system for providing the generated feedback to learners or customers.

[1276] A "generative AI model" is an artificial intelligence model that automatically generates optimal feedback and advice based on analysis results.

[1277] A "prompt" is a sentence that is input into a generative AI model and is an instruction sentence used to generate feedback or advice.

[1278] This invention is a system that can be applied not only to learners' learning activities but also to customer service support in brick-and-mortar stores. The system of the present invention is characterized by analyzing the behavior, facial expressions, and emotional states of learners and customers in real time and providing appropriate feedback and advice.

[1279] Hardware and Software

[1280] The main hardware used in this system is as follows:

[1281] Capture method: camera or smart glasses (e.g. Google Glass)

[1282] Transmission method: Tablet device or smartphone

[1283] Server: High-performance server (e.g. AWS EC2 instance)

[1284] The software used includes:

[1285] Video data analysis software: OpenCV

[1286] Sentiment analysis engine: Microsoft Azure Emotion API

[1287] Generative AI model: OpenAI's GPT-3 model

[1288] Data processing and calculation

[1289] The server performs the following steps:

[1290] 1. Acquisition of video data: Video data acquired by a camera is sent to a terminal. This data includes the behavior and facial expressions of the learner or customer.

[1291] 2. Data transmission: The device transmits the captured video data to the server, using data compression and a secure communication protocol.

[1292] 3. Video analysis and emotion assessment: The server analyzes the video data using OpenCV and assesses the emotional state using the Microsoft Azure Emotion API. From the analysis results, it identifies the learner's concentration level or fatigue level, and the customer's interest or confusion.

[1293] 4. Feedback generation: Based on the analysis results, the generative AI model (GPT-3) generates specific feedback and advice, including study advice and product information.

[1294] 5. Providing advice: The generated feedback is displayed and notified on the tablet device or smart glasses that are used as the means of providing the advice.

[1295] Examples and prompts

[1296] For example, in a brick-and-mortar scenario, consider a situation where a customer is interested in a particular product but is confused by its features. The server can analyze this confusion from the customer's facial expressions and behavior, and then use a generative AI model to input the following prompt:

[1297] "A customer was observed on camera. They were looking at product X with interest, but seemed confused about its features. Please generate additional information to provide to the customer."

[1298] Here, the generative AI model generates specific product descriptions and advice, which are displayed on the smart glasses, allowing store staff to provide information to customers at the right time.

[1299] In this way, a system can be built that provides optimal feedback to customers and learners in real time, achieving high-quality service and learning support.

[1300] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1301] Step 1:

[1302] The user begins learning activities or in-store activities. The imaging means (camera or smart glasses) captures the user's actions and facial expressions in real time. The input is the user's learning activities or in-store activities and facial expressions, and the output is video data.

[1303] Step 2:

[1304] The captured video data is sent to a device (tablet or smartphone). The device compresses the video data and sends it to a server using a secure communication protocol. The input is uncompressed video data, and the output is compressed video data.

[1305] Step 3:

[1306] The server analyzes the received video data. It processes visual information using OpenCV and evaluates emotional states using the Microsoft Azure Emotion API. From the analysis results, it identifies the user's behavior, facial expressions, concentration level, fatigue level, and customer interest or confusion. The input is compressed video data, and the output is analyzed behavioral and emotional data.

[1307] Step 4:

[1308] Based on the analysis results, the server uses a generative AI model (GPT-3) to generate specific feedback and advice. The prompt sentence is input into the generative AI model to obtain feedback appropriate for the specific situation. The input is the analyzed behavioral data, emotional data, and the prompt sentence, and the output is specific feedback and advice.

[1309] Step 5:

[1310] The means of providing this feedback (tablet device or smart glasses) displays the generated feedback and advice to the user, who then modifies or continues their learning or behavior accordingly. The input is the generated feedback and advice, and the output is the feedback provided to the user.

[1311] This process allows users to receive appropriate feedback in real time, enabling efficient learning and a meaningful purchasing experience.

[1312] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1313] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1314] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1315] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1316] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1317] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1318] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1319] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1320] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1321] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1322] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1323] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1324] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1325] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1326] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1327] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1328] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1329] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1330] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1331] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1332] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1333] The following is further disclosed regarding the above embodiment.

[1334] (Claim 1)

[1335] A photographing means for photographing the learning activities of learners;

[1336] a transmitting means for transmitting the video data captured by the imaging means to a server;

[1337] an analysis means in the server for analyzing the received video data;

[1338] a generation means for generating feedback appropriate for the learner based on the analysis result;

[1339] The system includes a means for providing the generated feedback to the learner.

[1340] (Claim 2)

[1341] 2. The system according to claim 1, wherein the imaging means detects the learner's facial expressions, movements, and line of sight, and estimates the learner's level of concentration and fatigue.

[1342] (Claim 3)

[1343] 2. The system according to claim 1, further comprising a learning progress management means for recording learning progress and adjusting the learning plan based on the saved progress data.

[1344] "Example 1"

[1345] (Claim 1)

[1346] A photographing means for photographing the learning activities of learners;

[1347] a transmitting means for compressing the video data captured by the capturing means and transmitting the compressed video data to a server using a secure communication protocol;

[1348] an analysis means in the server for analyzing the received video data using a specific software algorithm to analyze the learner's facial expressions, movements, and gaze;

[1349] a generation means for generating optimized feedback for a learner using a generative AI model based on the analysis results and past learning data;

[1350] The system includes a providing means for transmitting and displaying the generated feedback on the learner's mobile information terminal.

[1351] (Claim 2)

[1352] 2. The system according to claim 1, wherein the analysis means detects changes in the learner's facial expressions, movements, and line of sight, and evaluates the learner's level of concentration and fatigue.

[1353] (Claim 3)

[1354] 2. The system according to claim 1, further comprising a learning progress management means for recording the learning progress of the learner and adjusting the learning plan based on the stored progress data.

[1355] "Application Example 1"

[1356] (Claim 1)

[1357] A photographing means for photographing the learning activities of learners;

[1358] a transmitting means for transmitting the video data captured by the imaging means to a server;

[1359] an analysis means in the server for analyzing the received video data;

[1360] a generation means for generating feedback appropriate for the learner based on the analysis result;

[1361] a means for providing the generated feedback to the learner;

[1362] an automated equipment operation analysis means for analyzing the operation of automated equipment operating in a factory and providing feedback on improvements in efficiency and accuracy;

[1363] The system includes a display means for displaying the feedback.

[1364] (Claim 2)

[1365] 2. The system according to claim 1, wherein the imaging means detects the learner's facial expressions, movements, and gaze, estimates the learner's concentration level and fatigue level, or images and analyzes the operation of an automated device.

[1366] (Claim 3)

[1367] The system according to claim 1, further comprising a learning progress management means for recording learning progress and adjusting the learning plan based on the saved progress data, or for analyzing the operation of the automated device and providing feedback on areas for improvement.

[1368] "Example 2: Combining Emotion Engines"

[1369] (Claim 1)

[1370] A photographing means for photographing the learning activities of learners;

[1371] a transmitting means for transmitting the video data captured by the imaging means to a server;

[1372] an analysis means in the server for analyzing the received video data;

[1373] generating means for generating appropriate feedback based on the data and emotion data obtained by the analyzing means;

[1374] The system includes a means for providing the generated feedback to the learner.

[1375] (Claim 2)

[1376] 2. The system according to claim 1, wherein the imaging means detects the learner's facial expressions, movements, and line of sight, and estimates the learner's level of concentration and fatigue.

[1377] (Claim 3)

[1378] 10. The system of claim 1, further comprising emotion recognition means for analyzing the learner's emotional state from the video data and audio data.

[1379] (Claim 4)

[1380] 2. The system according to claim 1, further comprising a generating means for using a generative AI model to generate optimized feedback for the learner by integrating the acquired emotional data with data from the analyzing means.

[1381] (Claim 5)

[1382] 10. The system of claim 1, wherein the generated feedback includes specific advice, supplemental materials, additional questions, and words of encouragement.

[1383] (Claim 6)

[1384] 2. The system according to claim 1, further comprising a learning progress management means for recording learning progress and adjusting the learning plan based on the saved progress data.

[1385] "Application example 2 when combining emotion engines"

[1386] (Claim 1)

[1387] A photographing means for photographing the learning activities of learners;

[1388] a transmitting means for transmitting the video data captured by the imaging means to a server;

[1389] an analysis means in the server for analyzing the received video data;

[1390] a generation means for generating feedback appropriate for the learner based on the analysis result;

[1391] a providing means for providing the generated feedback to the learner; and

[1392] A means for capturing and evaluating the behavior and facial expressions of customers in the store;

[1393] A system that includes a generative AI model that generates optimal advice for customers based on the evaluation results.

[1394] (Claim 2)

[1395] The system according to claim 1, wherein the photographing means detects the facial expression, movement, and line of sight of the learner and estimates the degree of concentration and fatigue;

[1396] 10. The system of claim 1, further comprising means for assessing customer interest or confusion and generating appropriate prompt sentences.

[1397] (Claim 3)

[1398] The system according to claim 1, further comprising a learning progress management means for recording learning progress and adjusting the learning plan based on the saved progress data;

[1399] 10. The system of claim 1, further comprising means for providing appropriate advice in real time for customer service assistance. [Explanation of symbols]

[1400] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A photographing means for photographing the learning activities of learners; a transmitting means for transmitting the video data captured by the imaging means to a server; an analysis means in the server for analyzing the received video data; a generation means for generating feedback appropriate for the learner based on the analysis result; The system includes a means for providing the generated feedback to the learner.

2. 2. The system according to claim 1, wherein the image capturing means detects the facial expression, movement, and line of sight of the learner and estimates the degree of concentration and fatigue.

3. 2. The system according to claim 1, further comprising a learning progress management means for recording learning progress and adjusting the learning plan based on the saved progress data.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A