system
The system addresses limitations in conventional learning support by using optical character recognition and emotion analysis to provide flexible, personalized learning plans, optimizing learning experiences based on individual performance and emotional feedback.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-16
- Publication Date
- 2026-04-28
AI Technical Summary
Conventional learning support systems are limited to pre-registered problems and require external assistance for identifying learning strengths and weaknesses, lacking flexibility and efficiency in personalized learning support.
A system that utilizes optical character recognition to analyze user-captured images of test questions, classifies problems, and provides tailored learning plans based on individual performance, incorporating emotion analysis to adjust content difficulty and emotional feedback.
Enables efficient, personalized learning support that adapts to individual needs and emotional states, enhancing learning efficiency and effectiveness.
Smart Images

Figure 2026071021000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Conventional learning support systems are limited to problems registered in the platform in advance, so it is difficult to identify areas of strength and weakness from problems that do not exist in the platform, such as teaching materials and test questions. Also, there is a problem that in order to independently and efficiently grasp one's own learning situation, it is necessary to rely on a tutor or a cram school that involves time and cost.
Means for Solving the Problems
[0005] This invention flexibly handles questions outside of the platform by receiving images of test questions taken by the user with a smart device as input and using optical character recognition means to recognize characters from those images. As a result, when the user inputs their answer results, a problem analysis means identifies and classifies the problem, performs analysis in cooperation with a learning database, and then a learning presentation means presents a learning area suitable for the user based on that analysis. This enables the user to efficiently create a learning plan regardless of time or place.
[0006] "Image capture means" refers to a function that allows users to take pictures of exam questions using a digital device.
[0007] "Optical character recognition means" refers to a technology that extracts character information as digital data from captured images.
[0008] "Problem analysis means" refers to the process of identifying and classifying problems based on recognized character data.
[0009] "Learning presentation means" refers to a function that presents a learning area suitable for the user based on the results of the analyzed problem.
[0010] "Answer input method" refers to a function that allows users to input whether their answer is correct or incorrect.
[0011] A "server" refers to a computer system that provides data processing and analysis functions, and manages information by communicating with user terminals. [Brief explanation of the drawing]
[0012] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.
Embodiments for Carrying Out the Invention
[0013] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0014] First, the language used in the following description will be explained.
[0015] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0016] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0017] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs, various parameters, and the like. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.
[0018] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor, an antenna, and the like. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), and the like.
[0019] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0020] [First Embodiment]
[0021] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0022] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0023] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0024] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0025] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0026] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0027] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0028] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0029] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0030] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0031] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0032] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0033] The learning support system of the present invention operates by allowing users to capture exam questions and input answer information using a mobile device such as a smartphone or tablet. The user photographs the exam questions with a camera and indicates on the interface whether the answer is correct or incorrect. The captured image data and answer information are uploaded from the device to a server.
[0034] The server first performs optical character recognition (OCR) on the received image to extract text data. This process digitizes text information such as the question and answer choices. Next, the server analyzes the extracted text data, compares it with the registered learning database to identify and classify the question. Based on the identified question category and difficulty level, the server evaluates the user's answer results and analyzes areas that need to be studied.
[0035] The analysis results are sent from the server to the user's terminal. The terminal visually displays the received analysis results to the user. Based on these results, the user can identify their strengths and areas of learning that need improvement, and efficiently develop a learning plan.
[0036] For example, if a user takes a picture of a differential equation problem from high school mathematics and records an incorrect answer, the server categorizes the problem as a differential equation. It then provides the user with advice indicating that they need to strengthen their understanding of differential equations, from fundamentals to applications. In this way, the present invention provides efficient learning support tailored to the individual learning needs of each user.
[0037] The following describes the processing flow.
[0038] Step 1:
[0039] The user takes a picture of the test question using their smartphone camera. The user then reviews the image through an interface and selects whether their answer to the question is correct or incorrect.
[0040] Step 2:
[0041] The device packages the captured image and the user's entered answer information to send to the server, and uploads the data to the server via the network.
[0042] Step 3:
[0043] The server receives the uploaded image and uses OCR (Optical Character Recognition) technology to extract the text within the image. The extracted text includes the question and answer choices.
[0044] Step 4:
[0045] The server analyzes the text data obtained by OCR to identify and classify the problems. This includes matching the data against existing databases to determine the problem category and difficulty level.
[0046] Step 5:
[0047] The server considers the user's answer results (correct / incorrect) and analyzes the problem information to determine the optimal learning areas for the user. In particular, it identifies the user's weak areas and determines the areas that need to be strengthened intensively.
[0048] Step 6:
[0049] The server compiles the analysis results, creates a report in a format that is easy for the user to understand, and sends that data to the terminal.
[0050] Step 7:
[0051] The device receives analysis results sent from the server and displays them visually for the user. The user can then use this information to adjust their learning plan.
[0052] (Example 1)
[0053] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0054] There is a need for a system that can more efficiently identify the learning challenges each user has and provide optimal learning support tailored to their individual learning needs. However, conventional learning support systems have problems such as insufficient collection and analysis of user learning data, resulting in inadequate individual learning suggestions.
[0055] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0056] In this invention, the server includes means for capturing information using an information acquisition device, optical recognition means for recognizing data from the captured information, and information analysis means for identifying and classifying information based on the recognized data. This enables detailed analysis and optimized learning support tailored to individual learning situations.
[0057] An "information acquisition device" is a device used by users to capture data and information, and generally includes cameras and scanners.
[0058] An "optical recognition means" is a method for recognizing characters and data from captured information and extracting them as digital information, and optical character recognition (OCR) technology is commonly used.
[0059] "Information analysis means" refers to methods for identifying and classifying information based on recognized data and performing analyses related to user learning.
[0060] An "analysis presentation method" is a means of providing learners with visual or other feedback based on analyzed information, thereby supporting the optimization of their learning.
[0061] A "central processing unit" refers to a computing environment or server that operates in the background to receive and process data and generate analysis results.
[0062] The learning support system of the present invention operates by allowing the user to capture exam questions using a portable device, such as a smartphone or tablet, and input the answer information. The user uses the camera function of the portable device to take an image of the exam question. Next, the user inputs the answer through an interface within the device's application that specifies whether the answer is correct or incorrect.
[0063] The terminal transmits the captured image data and answer information to the server. The server uses optical character recognition software (e.g., Tesseract) to extract character data from the received image data. This enables the conversion of paper-based exam questions into digital data.
[0064] The server further analyzes the extracted text data, identifies similar information from the registered database, and classifies the problem. This analysis may utilize machine learning algorithms and database lookup techniques. After the analysis, the server evaluates the problem category and difficulty level, and analyzes which learning areas the user should strengthen. The final analysis results are sent to the user's terminal and provided as visual feedback.
[0065] For example, if a user takes a picture of a differential equation problem from high school mathematics and records their incorrect answer, the server analyzes the data and categorizes the problem as "differential equations." Based on the information obtained from the analysis, the server can provide the user with advice such as, "You need to strengthen your knowledge of differential equations, from basic to applied levels." This information helps the user to create an effective study plan.
[0066] An example of a prompt message would be, "Please suggest educational materials to deepen our understanding of the fundamentals of differential equations," which allows for further information gathering using a generative AI model.
[0067] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0068] Step 1:
[0069] The user uses the device to take a picture of the test questions with the camera. The image data is saved to the device as input. The user then inputs whether the answer is correct or incorrect on the device's interface. This information is associated with the captured image data, and the device prepares to send it to the server.
[0070] Step 2:
[0071] The device uploads the captured image data and associated answer information to the server. The image data and answer are transferred as input, and the server receives and stores this information. Secure protocols (such as HTTPS) are used for communication to ensure data security.
[0072] Step 3:
[0073] The server performs optical character recognition (OCR) processing on the received image data. The input is image data sent from the terminal, and the output is string data. Specifically, it uses OCR software (e.g., Tesseract) to convert the character information in the image into text.
[0074] Step 4:
[0075] The server analyzes the character data generated by OCR and compares it with a training database. The input is string data from OCR, and the output identifies the problem category and difficulty level. In this process, the server uses a text analysis algorithm to classify the problem.
[0076] Step 5:
[0077] The server evaluates the user's answers based on the analysis results. The input consists of categorized problem data and the user's answers, and the output proposes individual learning reinforcement areas. The server performs a statistical evaluation to analyze which areas the user needs to improve their knowledge in.
[0078] Step 6:
[0079] The server sends the analysis results to the user's terminal. The input is the analyzed data, and the output is specific learning advice for the user. The server structures this information and transfers it to the terminal.
[0080] Step 7:
[0081] The terminal visually presents the analysis results received from the server to the user. The input is the analysis data sent from the server, and the output is an information display that is intuitively understandable to the user. Specifically, it visualizes learning advice using graphs and charts.
[0082] (Application Example 1)
[0083] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0084] Early detection of abnormalities in equipment and components within a factory, and efficient identification of their causes, are crucial challenges in maintaining manufacturing lines. Conventional methods often involve delays in detecting abnormalities, leading to delayed responses. These challenges need to be addressed.
[0085] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0086] In this invention, the server includes an image capture means, a recognition means for recognizing information from the captured image, and an analysis means for identifying and classifying events based on the recognized information. This makes it possible to quickly detect abnormalities in equipment and parts within a factory, identify their causes, and prompt appropriate countermeasures.
[0087] "Image acquisition means" refers to the functions of devices or equipment used to capture images of a subject.
[0088] "Recognition means" refers to a function for extracting and identifying information from captured images.
[0089] "Analysis means" refers to a function for identifying and classifying events based on recognized information.
[0090] "Presentation means" refers to a function for providing the analyzed information to the user.
[0091] "Means for providing feedback through recognition and analysis means mounted on the device" refers to functions that provide rapid feedback using functions built into the device.
[0092] This system utilizes smart glasses and a cloud server. The smart glasses, with their built-in camera and interface, support on-site operations. Specifically, operators use the smart glasses to photograph abnormalities in machinery and equipment within the factory. This image data is transmitted to the cloud server via the internet. The server extracts text data from the images using OCR technology and further identifies and classifies abnormalities using AI analysis such as TENSORFLOW®. The analysis results are fed back to the smart glasses via the cloud, visually presenting the cause of the abnormality and countermeasures to the user. This enables rapid problem solving within the factory.
[0093] For example, if an operator detects an anomaly in a conveyor belt on a manufacturing line, they can use smart glasses to photograph the area. This image is sent to a cloud server, where AI analysis diagnoses that dirt has accumulated on the belt's sensor. The feedback includes specific instructions such as, "Please clean the conveyor belt sensor."
[0094] An example of a prompt message would be: "This image shows a part of a factory's conveyor line. If there is an abnormality, please tell me the cause and how to address it." Using this prompt message, the AI model derives appropriate analysis results.
[0095] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0096] Step 1:
[0097] The user uses smart glasses to photograph abnormal areas on machinery and equipment within the factory. The input is image data captured by the smart glasses' camera. This image data is temporarily stored in the terminal's memory for the next processing step.
[0098] Step 2:
[0099] The device uploads captured image data to a cloud server. The input is image data stored on the device, and the output is data sent to the server via the internet. This involves secure communication using security protocols such as SSL.
[0100] Step 3:
[0101] The server receives image data and uses OCR technology to extract text data. The input is image data, and the output is extracted text information. The OCR library (e.g., Tesseract) recognizes the text within the image.
[0102] Step 4:
[0103] The server uses OCR to extract text information, which is then analyzed using AI. The input is the text information obtained from the OCR, and the output is the result of identifying and classifying anomalies. Using generative AI models such as TensorFlow, the text information is analyzed to identify the anomalous parts and their causes.
[0104] Step 5:
[0105] The server feeds the analysis results back to the smart glasses terminal. The input is the result of the AI analysis, and the output is the feedback information sent to the user terminal. The data is transmitted over the network and displayed on the user interface.
[0106] Step 6:
[0107] The user receives feedback on anomalies through smart glasses and takes appropriate action. Based on the outputted information, the user implements specific countermeasures. For example, instructions such as "Please clean the conveyor belt sensor" are visually displayed.
[0108] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0109] The learning support system of this invention combines three main functions—image recognition, answer input, and emotion analysis—to maximize the user's learning efficiency. Users can use mobile devices such as smartphones and tablets to take pictures of exam questions and input their answers. In addition, the emotion engine analyzes the user's emotional state through the camera and microphone.
[0110] The terminal uploads images of the test questions and the user's answers to the server. At the same time, the user's facial expressions and voice data, captured by the device's camera and microphone, are also transmitted. The server first performs OCR on the images to extract text information. Next, a question analysis tool analyzes this text, comparing it to an existing database to identify and classify the questions.
[0111] Subsequently, the server uses an emotion engine to analyze the user's facial expressions and voice data to identify their emotional state. This emotion analysis is used to evaluate the user's stress level, concentration level, and other factors, and helps to adjust the learning presentation methods. Specifically, the difficulty and amount of learning content presented are adjusted based on the emotional state to optimize the learning experience and prevent the user from becoming overloaded.
[0112] For example, if the emotion engine detects signs of increased stress in the user when they select an incorrect answer to a difficult question, the server will adjust the difficulty level of the next learning question or display an encouraging message. In this way, the present invention enhances the quality of learning and provides personalized and appropriate learning support to the user.
[0113] The following describes the processing flow.
[0114] Step 1:
[0115] Users take photos of exam questions with their smartphone cameras and select whether their answers are correct or incorrect using an on-screen interface. The camera and microphone also capture the user's facial expressions and voice.
[0116] Step 2:
[0117] The device structures the data—including images, answer information, and acquired facial expressions and audio data—and uploads it to the server via the network in order to send it to the server.
[0118] Step 3:
[0119] The server receives the image data, first applying OCR technology to recognize the text within the image and extracting the character data of the question and answer choices.
[0120] Step 4:
[0121] The server performs problem analysis based on extracted text data, and identifies and classifies the problem category and difficulty level by referring to the platform's database.
[0122] Step 5:
[0123] The server uses an emotion engine to analyze received facial expression and voice data and evaluate the user's emotional state. This includes generating indicators such as stress levels and concentration levels.
[0124] Step 6:
[0125] The server combines the user's answer results with the sentiment analysis results and adjusts the learning content presented using a learning presentation method. Specifically, it optimizes the difficulty level and adds motivational messages.
[0126] Step 7:
[0127] The server compiles the adjusted learning suggestions into a report and sends it to the user's terminal.
[0128] Step 8:
[0129] The device receives analysis results from the server and presents them visually to the user. The user can then use this information to improve their learning plan.
[0130] (Example 2)
[0131] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0132] Conventional learning support systems fail to adjust learning content according to the user's emotional state, and in particular, they are insufficiently optimized for the learning experience in response to stress levels and concentration levels. Furthermore, they lack real-time feedback when users solve problems, making it difficult to improve learning efficiency.
[0133] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0134] In this invention, the server includes an image acquisition means, an optical character recognition means, an analysis means, an emotion analysis means, and an adjustment means. This makes it possible to extract and analyze character information from the user's image data, further analyze the user's emotional state, and dynamically adjust the learning content based on that.
[0135] "Image acquisition means" refers to a function that allows users to take pictures of learning materials using a mobile device.
[0136] "Optical character recognition means" refers to a technology that converts character information from image data into digital text.
[0137] "Analysis means" refers to a function that analyzes text information acquired by optical character recognition means, identifies the information, and classifies it.
[0138] "Presentation method" refers to a function that provides users with appropriate learning content based on analysis results and sentiment analysis results.
[0139] "Emotional analysis methods" refer to the process of analyzing a user's facial expressions and voice data to determine their emotional state.
[0140] The "adjustment mechanism" is a function that dynamically and appropriately changes the difficulty level and amount of learning content presented based on the emotion analysis results.
[0141] This learning support system operates through the collaboration of users, devices, and a server. Users can use devices such as smartphones and tablets to capture images of exam questions and input answer data during their studies. The device then sends the captured image data and answer data to the server.
[0142] The server extracts character information as digital text from received image data using high-precision optical character recognition (OCR) software. This technology can utilize existing OCR engines, such as Tesseract. The extracted text is analyzed by analysis software, which compares it against existing databases to identify and classify problems. Database management systems and natural language processing libraries are suitable for this analysis.
[0143] Next, the server supplies the user's facial expressions and voice data acquired through the camera and microphone to the emotion analysis engine to identify the user's emotional state. This analysis may utilize machine learning models, such as those using "OpenCV" or "Librosa."
[0144] The server adjusts the learning presentation methods based on the sentiment analysis results. Specifically, it changes the difficulty level and amount of learning content presented, and displays encouraging messages to the user, enabling them to continue learning efficiently. This optimizes the learning experience for each individual.
[0145] For example, if a user answers a difficult question incorrectly and the server detects an increase in stress through emotion analysis, it can maintain the user's motivation by making the next learning question easier or by providing an encouraging message.
[0146] An example of a prompt generated by a generative AI model is, "Please describe a system that adjusts learning content by considering the user's emotional state in order to maximize the user's learning efficiency."
[0147] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0148] Step 1:
[0149] Users take photos of exam questions using their mobile devices and input their answers. During this process, the device's camera and microphone simultaneously capture the user's facial expressions and voice. The input consists of the captured image, answer data, facial expression data, and voice data. This data is stored on the device and then prepared for transmission to the server.
[0150] Step 2:
[0151] The terminal sends image data, answer data, facial expression data, and audio data obtained from the user to the server. The input is the aforementioned dataset. For secure data transfer, it is desirable to use HTTPS as the protocol for transmission. As output, a receipt confirmation is displayed on the terminal.
[0152] Step 3:
[0153] The server processes the received image data into optical character recognition (OCR) software. The input is image data of the exam questions. OCR processing extracts the character information as digital text, which is then passed to the question analysis software. The output is the recognized character information.
[0154] Step 4:
[0155] The server analyzes text information, compares it with a database, and identifies and classifies the problems. The input is text information obtained by OCR. In this process, the type and difficulty level of the problems are identified, and they are classified as learning content. The output is the identified problem type and classification information.
[0156] Step 5:
[0157] The server passes the received facial expression and voice data to the emotion analysis engine. The input is biometric data. The emotion analysis engine processes the data to identify the user's emotional state, particularly stress levels and concentration levels. The output is information about the user's emotional state.
[0158] Step 6:
[0159] The server adjusts the presentation of learning content based on analyzed problem and emotional state information. Inputs are identified problem classifications and emotional states. Based on this data, the server adjusts the difficulty and amount of the next learning content presented, or generates encouraging messages. Outputs are the adjusted learning plan and feedback.
[0160] Step 7:
[0161] The server sends a customized learning plan and feedback to the user's device. The input consists of generated learning content and messages. Based on the received information, the device provides a learning experience optimized for the user. The output is the learning content and feedback displayed on the user's screen.
[0162] (Application Example 2)
[0163] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0164] In today's world, there is a demand for flexible learning support tailored to each user's individual learning pace and emotional state, but conventional systems have struggled to achieve this in real time and effectively. In particular, optimizing learning content while considering the user's emotional state is a challenge.
[0165] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0166] In this invention, the server includes an image acquisition device, a character recognition processing device, and a task analysis device. This enables dynamic adjustment of the learning content according to the user's answer information and emotional state.
[0167] An "image acquisition device" is a device that captures visual information in a specific environment, and includes cameras and scanners incorporated into digital devices.
[0168] A "character recognition processing device" is a processing device that uses optical methods to read character information from acquired image data and digitizes it.
[0169] A "task analysis device" is a device that has the function of identifying and classifying specific tasks or problems based on information obtained from a character recognition processing device.
[0170] A "learning display device" is a device that presents users with tailored learning content and has functions to promote an efficient learning experience.
[0171] An "emotion analysis device" is a device that analyzes a user's facial expressions and voice data to identify their emotional state.
[0172] An "adaptive learning device" is a device that uses the results of an emotion analysis device to dynamically adjust the difficulty level and amount of learning content to be optimal for the user.
[0173] To realize this invention, it is necessary to combine multiple hardware and software components. The system consists of a user's mobile terminal, a server, and various analysis programs.
[0174] The user's mobile device is equipped with a camera as an image acquisition device, which can be used to photograph learning materials and problems. This image data is transmitted to a server via a communication means. On the server, image recognition software is executed for optical character recognition, such as OpenCV. This software processes the extraction of character information from the image.
[0175] After character recognition is complete, the next component designed is a problem analysis unit. Here, the extracted text data is analyzed to identify and classify the problem. During this process, a database is used to match the problem with similar issues. Next, an emotion analysis unit analyzes the received user's facial expressions and voice data to determine their emotional state. Emotion analysis APIs such as Azure Cognitive Services are utilized for this purpose.
[0176] The results of the emotion analysis are sent to the adaptive learning device and used to adjust the learning content. For example, if the user is experiencing stress, the difficulty level of the task is adjusted. The content adjusted by the learning display device is then sent to the user's terminal, providing the user with an optimal learning experience.
[0177] As a concrete example, suppose a user is learning a mathematical formula. When they take a picture of the material with their camera, a simple question related to that formula is presented. Sentiment analysis is performed on the user's answer, and in some cases, an encouraging message is displayed.
[0178] An example of a prompt for a generative AI model is: "Consider how to adjust the learning content to match situations where the user feels burdened. Propose a method for analyzing the user's emotional state and appropriately adjusting the difficulty level."
[0179] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0180] Step 1:
[0181] The user takes a picture of the learning material using the camera on their mobile device. The input is an image of the learning material, and the output is image data. The device sends this image as digital data to the server.
[0182] Step 2:
[0183] The server processes the received image data using an optical character recognition (OCR) processing unit. The input is the transmitted image data, and the output is the extracted text information. The OpenCV library and other tools are used to extract the text information from the image.
[0184] Step 3:
[0185] The server analyzes the extracted text information using a problem analysis device. The input is text information, and the output is the identified problem and its classification result. The problem is identified by comparing it with similar known problems by referring to a database.
[0186] Step 4:
[0187] The user's facial expressions and voice data are transmitted from the terminal to the server. The input is the user's facial expressions and voice data, and the output is a state ready for emotion analysis. The data is converted to an appropriate format for use by the emotion analysis device.
[0188] Step 5:
[0189] The server uses an emotion analysis device to identify the user's emotional state. Input is facial expressions and voice data, and output is the user's emotional state. Azure Cognitive Services is used to analyze the user's stress and concentration levels.
[0190] Step 6:
[0191] The server adjusts the learning content using an adaptive learning device based on the emotion analysis results. The input is the emotional state and identified task, and the output is the adjusted learning content. The difficulty and amount of the learning content are dynamically adjusted to match the user's state.
[0192] Step 7:
[0193] The server sends the tailored learning content to the device. The input is the learning material, and the output is the educational content displayed to the user. The device uses a display device to visually present the content in order to provide a learning experience optimized for the user.
[0194] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0195] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0196] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0197] [Second Embodiment]
[0198] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0199] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0200] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0201] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0202] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0203] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0204] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0205] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0206] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0207] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0208] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0209] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0210] The learning support system of the present invention operates by allowing users to capture exam questions and input answer information using a mobile device such as a smartphone or tablet. The user photographs the exam questions with a camera and indicates on the interface whether the answer is correct or incorrect. The captured image data and answer information are uploaded from the device to a server.
[0211] The server first performs optical character recognition (OCR) on the received image to extract text data. This process digitizes text information such as the question and answer choices. Next, the server analyzes the extracted text data, compares it with the registered learning database to identify and classify the question. Based on the identified question category and difficulty level, the server evaluates the user's answer results and analyzes areas that need to be studied.
[0212] The analysis results are sent from the server to the user's terminal. The terminal visually displays the received analysis results to the user. Based on these results, the user can identify their strengths and areas of learning that need strengthening, and efficiently develop a learning plan.
[0213] For example, if a user takes a picture of a differential equation problem from high school mathematics and records an incorrect answer, the server categorizes the problem as a differential equation. It then provides the user with advice indicating that they need to strengthen their understanding of differential equations, from fundamentals to applications. In this way, the present invention provides efficient learning support tailored to the individual learning needs of each user.
[0214] The following describes the processing flow.
[0215] Step 1:
[0216] The user takes a picture of the exam question using their smartphone camera. The user then reviews the image through an interface and selects whether their answer to the question is correct or incorrect.
[0217] Step 2:
[0218] The device packages the captured image and the user's entered answer information to send to the server, and uploads the data to the server via the network.
[0219] Step 3:
[0220] The server receives the uploaded image and uses OCR (Optical Character Recognition) technology to extract the text within the image. The extracted text includes the question and answer choices.
[0221] Step 4:
[0222] The server analyzes the text data obtained by OCR to identify and classify the problems. This includes matching the data against existing databases to determine the problem category and difficulty level.
[0223] Step 5:
[0224] The server considers the user's answer results (correct / incorrect) and analyzes the problem information to determine the optimal learning areas for the user. In particular, it identifies the user's weak areas and determines areas that need focused reinforcement.
[0225] Step 6:
[0226] The server compiles the analysis results, creates a report in a format that is easy for the user to understand, and sends that data to the terminal.
[0227] Step 7:
[0228] The device receives analysis results sent from the server and displays them visually for the user. The user can then use this information to adjust their learning plan.
[0229] (Example 1)
[0230] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0231] There is a need for a system that can more efficiently identify the learning challenges each user has and provide optimal learning support tailored to their individual learning needs. However, conventional learning support systems have problems such as insufficient collection and analysis of user learning data, resulting in inadequate individual learning suggestions.
[0232] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0233] In this invention, the server includes means for capturing information using an information acquisition device, optical recognition means for recognizing data from the captured information, and information analysis means for identifying and classifying information based on the recognized data. This enables detailed analysis and optimized learning support tailored to individual learning situations.
[0234] An "information acquisition device" is a device used by users to capture data and information, and generally includes cameras and scanners.
[0235] An "optical recognition means" is a method for recognizing characters and data from captured information and extracting them as digital information, and optical character recognition (OCR) technology is commonly used.
[0236] "Information analysis means" refers to methods for identifying and classifying information based on recognized data and performing analysis related to user learning.
[0237] An "analysis presentation tool" is a means of providing learners with visual or other feedback based on analyzed information, thereby supporting the optimization of their learning.
[0238] A "central processing unit" refers to a computing environment or server that operates in the background to receive and process data and generate analysis results.
[0239] The learning support system of the present invention operates by allowing the user to capture exam questions using a portable device, such as a smartphone or tablet, and input the answer information. The user uses the camera function of the portable device to take an image of the exam question. Next, the user inputs the answer through an interface within the device's application that specifies whether the answer is correct or incorrect.
[0240] The terminal transmits the captured image data and answer information to the server. The server uses optical character recognition software (e.g., Tesseract) to extract character data from the received image data. This enables the conversion of paper-based exam questions into digital data.
[0241] The server further analyzes the extracted text data, identifies similar information from the registered database, and classifies the problem. This analysis may utilize machine learning algorithms and database lookup techniques. After the analysis, the server evaluates the problem category and difficulty level, and analyzes which learning areas the user should strengthen. The final analysis results are sent to the user's terminal and provided as visual feedback.
[0242] For example, if a user takes a picture of a differential equation problem from high school mathematics and records their incorrect answer, the server analyzes the data and categorizes the problem as "differential equations." Based on the information obtained from the analysis, the server can provide the user with advice such as, "You need to strengthen your knowledge of differential equations, from basic to applied levels." This information helps the user to create an effective study plan.
[0243] An example of a prompt message would be, "Please suggest educational materials to deepen our understanding of the fundamentals of differential equations," which allows for further information gathering using a generative AI model.
[0244] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0245] Step 1:
[0246] The user uses the device to take a picture of the test questions with the camera. The image data is saved to the device as input. The user then inputs whether the answer is correct or incorrect on the device's interface. This information is associated with the captured image data, and the device prepares to send it to the server.
[0247] Step 2:
[0248] The device uploads the captured image data and associated answer information to the server. The image data and answer are transferred as input, and the server receives and stores this information. Secure protocols (such as HTTPS) are used for communication to ensure data security.
[0249] Step 3:
[0250] The server performs optical character recognition (OCR) processing on the received image data. The input is image data sent from the terminal, and the output is string data. Specifically, it uses OCR software (e.g., Tesseract) to convert the character information in the image into text.
[0251] Step 4:
[0252] The server analyzes the character data generated by OCR and compares it with a training database. The input is string data from OCR, and the output identifies the problem category and difficulty level. In this process, the server uses a text analysis algorithm to classify the problem.
[0253] Step 5:
[0254] The server evaluates the user's answers based on the analysis results. The input consists of categorized problem data and the user's answers, and the output proposes individual learning reinforcement areas. The server performs a statistical evaluation to analyze which areas the user needs to improve their knowledge in.
[0255] Step 6:
[0256] The server sends the analysis results to the user's terminal. The input is the analyzed data, and the output is specific learning advice for the user. The server structures this information and transfers it to the terminal.
[0257] Step 7:
[0258] The terminal visually presents the analysis results received from the server to the user. The input is the analysis data sent from the server, and the output is an information display that is intuitively understandable to the user. Specifically, it visualizes learning advice using graphs and charts.
[0259] (Application Example 1)
[0260] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0261] Early detection of abnormalities in equipment and components within a factory, and efficient identification of their causes, are crucial challenges in maintaining manufacturing lines. Conventional methods often involve delays in detecting abnormalities, leading to delayed responses. These challenges need to be addressed.
[0262] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0263] In this invention, the server includes an image capture means, a recognition means for recognizing information from the captured image, and an analysis means for identifying and classifying events based on the recognized information. This makes it possible to quickly detect abnormalities in equipment and parts within a factory, identify their causes, and prompt appropriate countermeasures.
[0264] "Image acquisition means" refers to the functions of devices or equipment used to capture images of a subject.
[0265] "Recognition means" refers to a function for extracting and identifying information from captured images.
[0266] "Analysis means" refers to a function for identifying and classifying events based on recognized information.
[0267] "Presentation means" refers to a function for providing the analyzed information to the user.
[0268] "Means for providing feedback through recognition and analysis means mounted on the device" refers to functions that provide rapid feedback using functions built into the device.
[0269] This system utilizes smart glasses and a cloud server. The smart glasses, with their built-in camera and interface, support on-site operations. Specifically, operators use the smart glasses to photograph abnormalities in machinery and equipment within the factory. This image data is transmitted to the cloud server via the internet. The server extracts text data from the images using OCR technology and further identifies and classifies abnormalities using AI analysis such as TensorFlow. The analysis results are fed back to the smart glasses via the cloud, visually presenting the cause of the abnormality and countermeasures to the user. This enables rapid problem solving within the factory.
[0270] For example, if an operator detects an anomaly in a conveyor belt on a manufacturing line, they can use smart glasses to photograph the area. This image is sent to a cloud server, where AI analysis diagnoses that dirt has accumulated on the belt's sensor. The feedback includes specific instructions such as, "Please clean the conveyor belt sensor."
[0271] An example of a prompt message would be: "This image shows a part of a factory's conveyor line. If there is an abnormality, please tell me the cause and how to address it." Using this prompt message, the AI model derives appropriate analysis results.
[0272] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0273] Step 1:
[0274] The user uses smart glasses to photograph abnormal areas on machinery and equipment within the factory. The input is image data captured by the smart glasses' camera. This image data is temporarily stored in the terminal's memory for the next processing step.
[0275] Step 2:
[0276] The device uploads captured image data to a cloud server. The input is image data stored on the device, and the output is data sent to the server via the internet. This involves secure communication using security protocols such as SSL.
[0277] Step 3:
[0278] The server extracts text data from the received image data using OCR technology. The input is image data, and the output is extracted text information. The OCR library (e.g., Tesseract) recognizes the text within the image.
[0279] Step 4:
[0280] The server analyzes the character information extracted by OCR through AI analysis. The input is the character information obtained from OCR, and the output is the result of identifying and classifying abnormalities. Using a generative AI model such as TensorFlow, the text information is analyzed to identify abnormal parts and causes.
[0281] Step 5:
[0282] The server feeds back the analysis result to the smart glasses terminal. The input is the result of AI analysis, and the output is the feedback information sent to the user terminal. Data is transmitted through the network and displayed on the user interface.
[0283] Step 6:
[0284] The user receives feedback on the abnormality through the smart glasses and takes appropriate measures. Based on the output information, the user implements specific countermeasures. For example, instructions such as "Please clean the sensor of the conveyor belt" are visually displayed.
[0285] Furthermore, an emotion engine for estimating the user's emotion may be combined. That is, the specific processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform specific processing using the user's emotion.
[0286] The learning support system of the present invention combines three main functions of image recognition, answer input, and emotion analysis in order to maximize the user's learning efficiency. The user can use a mobile terminal such as a smartphone or tablet to take pictures of test questions and input answers. Also, through the camera and microphone, the emotion engine analyzes the user's emotional state.
[0287] The terminal uploads images of the test questions and the user's answers to the server. At the same time, the user's facial expressions and voice data, captured by the device's camera and microphone, are also transmitted. The server first performs OCR on the images to extract text information. Next, a question analysis tool analyzes this text, comparing it to an existing database to identify and classify the questions.
[0288] Subsequently, the server uses an emotion engine to analyze the user's facial expressions and voice data to identify their emotional state. This emotion analysis is used to evaluate the user's stress level, concentration level, and other factors, and helps to adjust the learning presentation methods. Specifically, the difficulty and amount of learning content presented are adjusted based on the emotional state to optimize the learning experience and prevent the user from becoming overloaded.
[0289] For example, if the emotion engine detects signs of increased stress in the user when they select an incorrect answer to a difficult question, the server will adjust the difficulty level of the next learning question or display an encouraging message. In this way, the present invention enhances the quality of learning and provides personalized and appropriate learning support to the user.
[0290] The following describes the processing flow.
[0291] Step 1:
[0292] Users take photos of exam questions with their smartphone cameras and select whether their answers are correct or incorrect using an on-screen interface. The camera and microphone also capture the user's facial expressions and voice.
[0293] Step 2:
[0294] The device structures the data—including images, answer information, and acquired facial expressions and audio data—and uploads it to the server via the network in order to send it to the server.
[0295] Step 3:
[0296] The server receives the image data, first applies OCR technology to recognize the text in the image, and extracts the character data of the question text and answer options.
[0297] Step 4:
[0298] Based on the character data extracted by the server, the server performs question analysis, refers to the database of the platform to identify and classify the category and difficulty level of the question.
[0299] Step 5:
[0300] The server uses the emotion engine to analyze the received facial expression data and voice data, and evaluates the user's emotional state. This includes generating indicators indicating the stress level and concentration.
[0301] Step 6:
[0302] The server combines the user's answer result and the result of emotion analysis, and adjusts the learning content to be presented using the learning prompting means. Specifically, it optimizes the difficulty level and adds motivation messages.
[0303] Step 7:
[0304] The server summarizes the adjusted learning proposal as a report and transmits it to the user's terminal.
[0305] Step 8:
[0306] The terminal receives the analysis result from the server and visually presents it to the user. The user can refer to this to improve their learning plan.
[0307] (Example 2)
[0308] Next, Example 2 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".
[0309] Conventional learning support systems fail to adjust learning content according to the user's emotional state, and in particular, they are insufficiently optimized for the learning experience in response to stress levels and concentration levels. Furthermore, they lack real-time feedback when users solve problems, making it difficult to improve learning efficiency.
[0310] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0311] In this invention, the server includes an image acquisition means, an optical character recognition means, an analysis means, an emotion analysis means, and an adjustment means. This makes it possible to extract and analyze character information from the user's image data, further analyze the user's emotional state, and dynamically adjust the learning content based on that.
[0312] "Image acquisition means" refers to a function that allows users to take pictures of learning materials using a mobile device.
[0313] "Optical character recognition means" refers to a technology that converts character information from image data into digital text.
[0314] "Analysis means" refers to a function that analyzes text information acquired by optical character recognition means, identifies the information, and classifies it.
[0315] "Presentation method" refers to a function that provides users with appropriate learning content based on analysis results and sentiment analysis results.
[0316] "Emotional analysis methods" refer to the process of analyzing a user's facial expressions and voice data to determine their emotional state.
[0317] The "adjustment mechanism" is a function that dynamically and appropriately changes the difficulty level and amount of learning content presented based on the emotion analysis results.
[0318] This learning support system operates through the collaboration of users, devices, and a server. Users can use devices such as smartphones and tablets to capture images of exam questions and input answer data during their studies. The device then sends the captured image data and answer data to the server.
[0319] The server extracts character information as digital text from received image data using high-precision optical character recognition (OCR) software. This technology can utilize existing OCR engines, such as Tesseract. The extracted text is analyzed by analysis software, which compares it against existing databases to identify and classify problems. Database management systems and natural language processing libraries are suitable for this analysis.
[0320] Next, the server supplies the user's facial expressions and voice data acquired through the camera and microphone to the emotion analysis engine to identify the user's emotional state. This analysis may utilize machine learning models, such as those using "OpenCV" or "Librosa."
[0321] The server adjusts the learning presentation methods based on the sentiment analysis results. Specifically, it changes the difficulty level and amount of learning content presented, and displays encouraging messages to the user, enabling them to continue learning efficiently. This optimizes the learning experience for each individual.
[0322] For example, if a user answers a difficult question incorrectly and the server detects an increase in stress through emotion analysis, it can maintain the user's motivation by making the next learning question easier or by providing an encouraging message.
[0323] An example of a prompt generated by a generative AI model is, "Please describe a system that adjusts learning content by considering the user's emotional state in order to maximize the user's learning efficiency."
[0324] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0325] Step 1:
[0326] Users take photos of exam questions using their mobile devices and input their answers. During this process, the device's camera and microphone simultaneously capture the user's facial expressions and voice. The input consists of the captured image, answer data, facial expression data, and voice data. This data is stored on the device and then prepared for transmission to the server.
[0327] Step 2:
[0328] The terminal sends image data, answer data, facial expression data, and audio data obtained from the user to the server. The input is the aforementioned dataset. For secure data transfer, it is desirable to use HTTPS as the protocol for transmission. As output, a receipt confirmation is displayed on the terminal.
[0329] Step 3:
[0330] The server processes the received image data into optical character recognition (OCR) software. The input is image data of the exam questions. OCR processing extracts the character information as digital text, which is then passed to the question analysis software. The output is the recognized character information.
[0331] Step 4:
[0332] The server analyzes text information, compares it with a database, and identifies and classifies the problems. The input is text information obtained by OCR. In this process, the type and difficulty level of the problems are identified, and they are classified as learning content. The output is the identified problem type and classification information.
[0333] Step 5:
[0334] The server passes the received facial expression and voice data to the emotion analysis engine. The input is biometric data. The emotion analysis engine processes the data to identify the user's emotional state, particularly stress levels and concentration levels. The output is information about the user's emotional state.
[0335] Step 6:
[0336] The server adjusts the presentation of learning content based on analyzed problem and emotional state information. Inputs are identified problem classifications and emotional states. Based on this data, the server adjusts the difficulty and amount of the next learning content presented, or generates encouraging messages. Outputs are the adjusted learning plan and feedback.
[0337] Step 7:
[0338] The server sends a customized learning plan and feedback to the user's device. The input consists of generated learning content and messages. Based on the received information, the device provides a learning experience optimized for the user. The output is the learning content and feedback displayed on the user's screen.
[0339] (Application Example 2)
[0340] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the smart glasses 214 as the "terminal".
[0341] In today's world, there is a demand for flexible learning support tailored to each user's individual learning pace and emotional state, but conventional systems have struggled to achieve this in real time and effectively. In particular, optimizing learning content while considering the user's emotional state is a challenge.
[0342] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0343] In this invention, the server includes an image acquisition device, a character recognition processing device, and a task analysis device. This enables dynamic adjustment of the learning content according to the user's answer information and emotional state.
[0344] An "image acquisition device" is a device that captures visual information in a specific environment, and includes cameras and scanners incorporated into digital devices.
[0345] A "character recognition processing device" is a processing device that uses optical methods to read character information from acquired image data and digitizes it.
[0346] A "task analysis device" is a device that has the function of identifying and classifying specific tasks or problems based on information obtained from a character recognition processing device.
[0347] A "learning display device" is a device that presents users with tailored learning content and has functions to promote an efficient learning experience.
[0348] An "emotion analysis device" is a device that analyzes a user's facial expressions and voice data to identify their emotional state.
[0349] An "adaptive learning device" is a device that uses the results of an emotion analysis device to dynamically adjust the difficulty level and amount of learning content to be optimal for the user.
[0350] To realize this invention, it is necessary to combine multiple hardware and software components. The system consists of a user's mobile terminal, a server, and various analysis programs.
[0351] The user's mobile device is equipped with a camera as an image acquisition device, which can be used to photograph learning materials and problems. This image data is transmitted to a server via a communication means. On the server, image recognition software is executed for optical character recognition, such as OpenCV. This software processes the extraction of character information from the image.
[0352] After character recognition is complete, the next component designed is a problem analysis unit. Here, the extracted text data is analyzed to identify and classify the problem. A database is used in this process to match the problem with similar issues. Next, an emotion analysis unit analyzes the received user's facial expressions and voice data to determine their emotional state. Emotion analysis APIs such as Azure Cognitive Services are utilized for this purpose.
[0353] The results of the emotion analysis are sent to the adaptive learning device and used to adjust the learning content. For example, if the user is experiencing stress, the difficulty level of the task is adjusted. The content adjusted by the learning display device is then sent to the user's terminal, providing the user with an optimal learning experience.
[0354] As a concrete example, suppose a user is learning a mathematical formula. When they take a picture of the material with their camera, a simple question related to that formula is presented. Sentiment analysis is performed on the user's answer, and in some cases, an encouraging message is displayed.
[0355] An example of a prompt for a generative AI model is: "Consider how to adjust the learning content to match situations where the user feels burdened. Propose a method for analyzing the user's emotional state and appropriately adjusting the difficulty level."
[0356] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0357] Step 1:
[0358] The user takes a picture of the learning material using the camera on their mobile device. The input is an image of the learning material, and the output is image data. The device sends this image as digital data to the server.
[0359] Step 2:
[0360] The server processes the received image data using an optical character recognition (OCR) processing unit. The input is the transmitted image data, and the output is the extracted text information. The OpenCV library and other tools are used to extract the text information from the image.
[0361] Step 3:
[0362] The server analyzes the extracted text information using a problem analysis device. The input is text information, and the output is the identified problem and its classification result. The problem is identified by comparing it with similar known problems by referring to a database.
[0363] Step 4:
[0364] The user's facial expressions and voice data are transmitted from the terminal to the server. The input is the user's facial expressions and voice data, and the output is a state ready for emotion analysis. The data is converted to an appropriate format for use by the emotion analysis device.
[0365] Step 5:
[0366] The server uses an emotion analysis device to identify the user's emotional state. Input is facial expressions and voice data, and output is the user's emotional state. Azure Cognitive Services is used to analyze the user's stress and concentration levels.
[0367] Step 6:
[0368] The server adjusts the learning content using an adaptive learning device based on the emotion analysis results. The input is the emotional state and identified task, and the output is the adjusted learning content. The difficulty and amount of the learning content are dynamically adjusted to match the user's state.
[0369] Step 7:
[0370] The server sends the tailored learning content to the device. The input is the learning material, and the output is the educational content displayed to the user. The device uses a display device to visually present the content in order to provide a learning experience optimized for the user.
[0371] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0372] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0373] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0374] [Third Embodiment]
[0375] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0376] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0377] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0378] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0379] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0380] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0381] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0382] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0383] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0384] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0385] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0386] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0387] The learning support system of the present invention operates by allowing users to capture exam questions and input answer information using a mobile device such as a smartphone or tablet. The user photographs the exam questions with a camera and indicates on the interface whether the answer is correct or incorrect. The captured image data and answer information are uploaded from the device to a server.
[0388] The server first performs optical character recognition (OCR) on the received image to extract text data. This process digitizes text information such as the question and answer choices. Next, the server analyzes the extracted text data, compares it with the registered learning database to identify and classify the question. Based on the identified question category and difficulty level, the server evaluates the user's answer results and analyzes areas that need to be studied.
[0389] The analysis results are sent from the server to the user's terminal. The terminal visually displays the received analysis results to the user. Based on these results, the user can identify their strengths and areas of learning that need strengthening, and efficiently develop a learning plan.
[0390] For example, if a user takes a picture of a differential equation problem from high school mathematics and records an incorrect answer, the server categorizes the problem as a differential equation. It then provides the user with advice indicating that they need to strengthen their understanding of differential equations, from fundamentals to applications. In this way, the present invention provides efficient learning support tailored to the individual learning needs of each user.
[0391] The following describes the processing flow.
[0392] Step 1:
[0393] The user takes a picture of the exam question using their smartphone camera. The user then reviews the image through an interface and selects whether their answer to the question is correct or incorrect.
[0394] Step 2:
[0395] The device packages the captured image and the user's entered answer information to send to the server, and uploads the data to the server via the network.
[0396] Step 3:
[0397] The server receives the uploaded image and uses OCR (Optical Character Recognition) technology to extract the text within the image. The extracted text includes the question and answer choices.
[0398] Step 4:
[0399] The server analyzes the text data obtained by OCR to identify and classify the problems. This includes matching the data against existing databases to determine the problem category and difficulty level.
[0400] Step 5:
[0401] The server considers the user's answer results (correct / incorrect) and analyzes the problem information to determine the optimal learning areas for the user. In particular, it identifies the user's weak areas and determines areas that need focused reinforcement.
[0402] Step 6:
[0403] The server compiles the analysis results, creates a report in a format that is easy for the user to understand, and sends that data to the terminal.
[0404] Step 7:
[0405] The device receives analysis results sent from the server and displays them visually for the user. The user can then use this information to adjust their learning plan.
[0406] (Example 1)
[0407] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0408] There is a need for a system that can more efficiently identify the learning challenges each user has and provide optimal learning support tailored to their individual learning needs. However, conventional learning support systems have problems such as insufficient collection and analysis of user learning data, resulting in inadequate individual learning suggestions.
[0409] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0410] In this invention, the server includes means for capturing information using an information acquisition device, optical recognition means for recognizing data from the captured information, and information analysis means for identifying and classifying information based on the recognized data. This enables detailed analysis and optimized learning support tailored to individual learning situations.
[0411] An "information acquisition device" is a device used by users to capture data and information, and generally includes cameras and scanners.
[0412] An "optical recognition means" is a method for recognizing characters and data from captured information and extracting them as digital information, and optical character recognition (OCR) technology is commonly used.
[0413] "Information analysis means" refers to methods for identifying and classifying information based on recognized data and performing analysis related to user learning.
[0414] An "analysis presentation tool" is a means of providing learners with visual or other feedback based on analyzed information, thereby supporting the optimization of their learning.
[0415] A "central processing unit" refers to a computing environment or server that operates in the background to receive and process data and generate analysis results.
[0416] The learning support system of the present invention operates by allowing the user to capture exam questions using a portable device, such as a smartphone or tablet, and input the answer information. The user uses the camera function of the portable device to take an image of the exam question. Next, the user inputs the answer through an interface within the device's application that specifies whether the answer is correct or incorrect.
[0417] The terminal transmits the captured image data and answer information to the server. The server uses optical character recognition software (e.g., Tesseract) to extract character data from the received image data. This enables the conversion of paper-based exam questions into digital data.
[0418] The server further analyzes the extracted text data, identifies similar information from the registered database, and classifies the problem. This analysis may utilize machine learning algorithms and database lookup techniques. After the analysis, the server evaluates the problem category and difficulty level, and analyzes which learning areas the user should strengthen. The final analysis results are sent to the user's terminal and provided as visual feedback.
[0419] For example, if a user takes a picture of a differential equation problem from high school mathematics and records their incorrect answer, the server analyzes the data and categorizes the problem as "differential equations." Based on the information obtained from the analysis, the server can provide the user with advice such as, "You need to strengthen your knowledge of differential equations, from basic to applied levels." This information helps the user to create an effective study plan.
[0420] An example of a prompt message would be, "Please suggest educational materials to deepen our understanding of the fundamentals of differential equations," which allows for further information gathering using a generative AI model.
[0421] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0422] Step 1:
[0423] The user uses the device to take a picture of the test questions with the camera. The image data is saved to the device as input. The user then inputs whether the answer is correct or incorrect on the device's interface. This information is associated with the captured image data, and the device prepares to send it to the server.
[0424] Step 2:
[0425] The device uploads the captured image data and associated answer information to the server. The image data and answer are transferred as input, and the server receives and stores this information. Secure protocols (such as HTTPS) are used for communication to ensure data security.
[0426] Step 3:
[0427] The server performs optical character recognition (OCR) processing on the received image data. The input is image data sent from the terminal, and the output is string data. Specifically, it uses OCR software (e.g., Tesseract) to convert the character information in the image into text.
[0428] Step 4:
[0429] The server analyzes the character data generated by OCR and compares it with a training database. The input is string data from OCR, and the output identifies the problem category and difficulty level. In this process, the server uses a text analysis algorithm to classify the problem.
[0430] Step 5:
[0431] The server evaluates the user's answers based on the analysis results. The input consists of categorized problem data and the user's answers, and the output proposes individual learning reinforcement areas. The server performs a statistical evaluation to analyze which areas the user needs to improve their knowledge in.
[0432] Step 6:
[0433] The server sends the analysis results to the user's terminal. The input is the analyzed data, and the output is specific learning advice for the user. The server structures this information and transfers it to the terminal.
[0434] Step 7:
[0435] The terminal visually presents the analysis results received from the server to the user. The input is the analysis data sent from the server, and the output is an information display that is intuitively understandable to the user. Specifically, it visualizes learning advice using graphs and charts.
[0436] (Application Example 1)
[0437] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0438] Early detection of abnormalities in equipment and components within a factory, and efficient identification of their causes, are crucial challenges in maintaining manufacturing lines. Conventional methods often involve delays in detecting abnormalities, leading to delayed responses. These challenges need to be addressed.
[0439] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0440] In this invention, the server includes an image capture means, a recognition means for recognizing information from the captured image, and an analysis means for identifying and classifying events based on the recognized information. This makes it possible to quickly detect abnormalities in equipment and parts within a factory, identify their causes, and prompt appropriate countermeasures.
[0441] "Image acquisition means" refers to the functions of devices or equipment used to capture images of a subject.
[0442] "Recognition means" refers to a function for extracting and identifying information from captured images.
[0443] "Analysis means" refers to a function for identifying and classifying events based on recognized information.
[0444] "Presentation means" refers to a function for providing the analyzed information to the user.
[0445] "Means for providing feedback through recognition and analysis means mounted on the device" refers to functions that provide rapid feedback using functions built into the device.
[0446] This system utilizes smart glasses and a cloud server. The smart glasses, with their built-in camera and interface, support on-site operations. Specifically, operators use the smart glasses to photograph abnormalities in machinery and equipment within the factory. This image data is transmitted to the cloud server via the internet. The server extracts text data from the images using OCR technology and further identifies and classifies abnormalities using AI analysis such as TensorFlow. The analysis results are fed back to the smart glasses via the cloud, visually presenting the cause of the abnormality and countermeasures to the user. This enables rapid problem solving within the factory.
[0447] For example, if an operator detects an anomaly in a conveyor belt on a manufacturing line, they can use smart glasses to photograph the area. This image is sent to a cloud server, where AI analysis diagnoses that dirt has accumulated on the belt's sensor. The feedback includes specific instructions such as, "Please clean the conveyor belt sensor."
[0448] An example of a prompt message would be: "This image shows a part of a factory's conveyor line. If there is an abnormality, please tell me the cause and how to address it." Using this prompt message, the AI model derives appropriate analysis results.
[0449] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0450] Step 1:
[0451] The user uses smart glasses to photograph abnormal areas on machinery and equipment within the factory. The input is image data captured by the smart glasses' camera. This image data is temporarily stored in the terminal's memory for the next processing step.
[0452] Step 2:
[0453] The device uploads captured image data to a cloud server. The input is image data stored on the device, and the output is data sent to the server via the internet. This involves secure communication using security protocols such as SSL.
[0454] Step 3:
[0455] The server extracts text data from the received image data using OCR technology. The input is image data, and the output is extracted text information. The OCR library (e.g., Tesseract) recognizes the text within the image.
[0456] Step 4:
[0457] The server uses OCR to extract text information, which is then analyzed using AI. The input is the text information obtained from OCR, and the output is the result of identifying and classifying anomalies. Using generative AI models such as TensorFlow, the text information is analyzed to identify the anomalous parts and their causes.
[0458] Step 5:
[0459] The server feeds the analysis results back to the smart glasses terminal. The input is the result of the AI analysis, and the output is the feedback information sent to the user terminal. The data is transmitted over the network and displayed on the user interface.
[0460] Step 6:
[0461] The user receives feedback on anomalies through smart glasses and takes appropriate action. Based on the outputted information, the user implements specific countermeasures. For example, instructions such as "Please clean the conveyor belt sensor" are visually displayed.
[0462] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0463] The learning support system of this invention combines three main functions—image recognition, answer input, and emotion analysis—to maximize the user's learning efficiency. Users can use mobile devices such as smartphones and tablets to take pictures of exam questions and input their answers. In addition, the emotion engine analyzes the user's emotional state through the camera and microphone.
[0464] The terminal uploads images of the test questions and the user's answers to the server. At the same time, the user's facial expressions and voice data, captured by the device's camera and microphone, are also transmitted. The server first performs OCR on the images to extract text information. Next, a question analysis tool analyzes this text, comparing it to an existing database to identify and classify the questions.
[0465] Subsequently, the server uses an emotion engine to analyze the user's facial expressions and voice data to identify their emotional state. This emotion analysis is used to evaluate the user's stress level, concentration level, and other factors, and helps to adjust the learning presentation methods. Specifically, the difficulty and amount of learning content presented are adjusted based on the emotional state to optimize the learning experience and prevent the user from becoming overloaded.
[0466] For example, if the emotion engine detects signs of increased stress in the user when they select an incorrect answer to a difficult question, the server will adjust the difficulty level of the next learning question or display an encouraging message. In this way, the present invention enhances the quality of learning and provides personalized and appropriate learning support to the user.
[0467] The following describes the processing flow.
[0468] Step 1:
[0469] Users take photos of exam questions with their smartphone cameras and select whether their answers are correct or incorrect using an on-screen interface. The camera and microphone also capture the user's facial expressions and voice.
[0470] Step 2:
[0471] The device structures the data—including images, answer information, and acquired facial expressions and audio data—and uploads it to the server via the network in order to send it to the server.
[0472] Step 3:
[0473] The server receives the image data, first applying OCR technology to recognize the text within the image and extracting the character data of the question and answer choices.
[0474] Step 4:
[0475] The server performs problem analysis based on extracted text data, and identifies and classifies the problem category and difficulty level by referring to the platform's database.
[0476] Step 5:
[0477] The server uses an emotion engine to analyze received facial expression and voice data and evaluate the user's emotional state. This includes generating indicators such as stress levels and concentration levels.
[0478] Step 6:
[0479] The server combines the user's answer results with the sentiment analysis results and adjusts the learning content presented using a learning presentation method. Specifically, it optimizes the difficulty level and adds motivational messages.
[0480] Step 7:
[0481] The server compiles the adjusted learning suggestions into a report and sends it to the user's terminal.
[0482] Step 8:
[0483] The device receives analysis results from the server and presents them visually to the user. The user can then use this information to improve their learning plan.
[0484] (Example 2)
[0485] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0486] Conventional learning support systems fail to adjust learning content according to the user's emotional state, and in particular, they are insufficiently optimized for the learning experience in response to stress levels and concentration levels. Furthermore, they lack real-time feedback when users solve problems, making it difficult to improve learning efficiency.
[0487] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0488] In this invention, the server includes an image acquisition means, an optical character recognition means, an analysis means, an emotion analysis means, and an adjustment means. This makes it possible to extract and analyze character information from the user's image data, further analyze the user's emotional state, and dynamically adjust the learning content based on that.
[0489] "Image acquisition means" refers to a function that allows users to take pictures of learning materials using a mobile device.
[0490] "Optical character recognition means" refers to a technology that converts character information from image data into digital text.
[0491] "Analysis means" refers to a function that analyzes text information acquired by optical character recognition means, identifies the information, and classifies it.
[0492] "Presentation method" refers to a function that provides users with appropriate learning content based on analysis results and sentiment analysis results.
[0493] "Emotional analysis methods" refer to the process of analyzing a user's facial expressions and voice data to determine their emotional state.
[0494] The "adjustment mechanism" is a function that dynamically and appropriately changes the difficulty level and amount of learning content presented based on the emotion analysis results.
[0495] This learning support system operates through the collaboration of users, devices, and a server. Users can use devices such as smartphones and tablets to capture images of exam questions and input answer data during their studies. The device then sends the captured image data and answer data to the server.
[0496] The server extracts character information as digital text from received image data using high-precision optical character recognition (OCR) software. This technology can utilize existing OCR engines, such as Tesseract. The extracted text is analyzed by analysis software, which compares it against existing databases to identify and classify problems. Database management systems and natural language processing libraries are suitable for this analysis.
[0497] Next, the server supplies the user's facial expressions and voice data acquired through the camera and microphone to the emotion analysis engine to identify the user's emotional state. This analysis may utilize machine learning models, such as those using "OpenCV" or "Librosa."
[0498] The server adjusts the learning presentation methods based on the sentiment analysis results. Specifically, it changes the difficulty level and amount of learning content presented, and displays encouraging messages to the user, enabling them to continue learning efficiently. This optimizes the learning experience for each individual.
[0499] For example, if a user answers a difficult question incorrectly and the server detects an increase in stress through emotion analysis, it can maintain the user's motivation by making the next learning question easier or by providing an encouraging message.
[0500] An example of a prompt generated by a generative AI model is, "Please describe a system that adjusts learning content by considering the user's emotional state in order to maximize the user's learning efficiency."
[0501] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0502] Step 1:
[0503] Users take photos of exam questions using their mobile devices and input their answers. During this process, the device's camera and microphone simultaneously capture the user's facial expressions and voice. The input consists of the captured image, answer data, facial expression data, and voice data. This data is stored on the device and then prepared for transmission to the server.
[0504] Step 2:
[0505] The terminal sends image data, answer data, facial expression data, and audio data obtained from the user to the server. The input is the aforementioned dataset. For secure data transfer, it is desirable to use HTTPS as the protocol for transmission. As output, a receipt confirmation is displayed on the terminal.
[0506] Step 3:
[0507] The server processes the received image data into optical character recognition (OCR) software. The input is image data of the exam questions. The OCR process extracts the character information as digital text, which is then passed to the question analysis software. The output is the recognized character information.
[0508] Step 4:
[0509] The server analyzes text information, compares it with a database, and identifies and classifies the problems. The input is text information obtained by OCR. In this process, the type and difficulty level of the problems are identified, and they are classified as learning content. The output is the identified problem type and classification information.
[0510] Step 5:
[0511] The server passes the received facial expression and voice data to the emotion analysis engine. The input is biometric data. The emotion analysis engine processes the data to identify the user's emotional state, particularly stress levels and concentration levels. The output is information about the user's emotional state.
[0512] Step 6:
[0513] The server adjusts the presentation of learning content based on analyzed problem and emotional state information. Inputs are identified problem classifications and emotional states. Based on this data, the server adjusts the difficulty and amount of the next learning content presented, or generates encouraging messages. Outputs are the adjusted learning plan and feedback.
[0514] Step 7:
[0515] The server sends a customized learning plan and feedback to the user's device. The input consists of generated learning content and messages. Based on the received information, the device provides a learning experience optimized for the user. The output is the learning content and feedback displayed on the user's screen.
[0516] (Application Example 2)
[0517] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0518] In today's world, there is a demand for flexible learning support tailored to each user's individual learning pace and emotional state, but conventional systems have struggled to achieve this in real time and effectively. In particular, optimizing learning content while considering the user's emotional state is a challenge.
[0519] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0520] In this invention, the server includes an image acquisition device, a character recognition processing device, and a task analysis device. This enables dynamic adjustment of the learning content according to the user's answer information and emotional state.
[0521] An "image acquisition device" is a device that captures visual information in a specific environment, and includes cameras and scanners incorporated into digital devices.
[0522] A "character recognition processing device" is a processing device that uses optical methods to read character information from acquired image data and digitizes it.
[0523] A "task analysis device" is a device that has the function of identifying and classifying specific tasks or problems based on information obtained from a character recognition processing device.
[0524] A "learning display device" is a device that presents users with tailored learning content and has functions to promote an efficient learning experience.
[0525] An "emotion analysis device" is a device that analyzes a user's facial expressions and voice data to identify their emotional state.
[0526] An "adaptive learning device" is a device that uses the results of an emotion analysis device to dynamically adjust the difficulty level and amount of learning content to be optimal for the user.
[0527] To realize this invention, it is necessary to combine multiple hardware and software components. The system consists of a user's mobile terminal, a server, and various analysis programs.
[0528] The user's mobile device is equipped with a camera as an image acquisition device, which can be used to photograph learning materials and problems. This image data is transmitted to a server via a communication means. On the server, image recognition software is executed for optical character recognition, such as OpenCV. This software processes the extraction of character information from the image.
[0529] After character recognition is complete, the next component designed is a problem analysis unit. Here, the extracted text data is analyzed to identify and classify the problem. A database is used in this process to match the problem with similar issues. Next, an emotion analysis unit analyzes the received user's facial expressions and voice data to determine their emotional state. Emotion analysis APIs such as Azure Cognitive Services are utilized for this purpose.
[0530] The results of the emotion analysis are sent to the adaptive learning device and used to adjust the learning content. For example, if the user is experiencing stress, the difficulty level of the task is adjusted. The content adjusted by the learning display device is then sent to the user's terminal, providing the user with an optimal learning experience.
[0531] As a concrete example, suppose a user is learning a mathematical formula. When they take a picture of the material with their camera, a simple question related to that formula is presented. Sentiment analysis is performed on the user's answer, and in some cases, an encouraging message is displayed.
[0532] An example of a prompt for a generative AI model is: "Consider how to adjust the learning content to match situations where the user feels burdened. Propose a method for analyzing the user's emotional state and appropriately adjusting the difficulty level."
[0533] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0534] Step 1:
[0535] The user takes a picture of the learning material using the camera on their mobile device. The input is an image of the learning material, and the output is image data. The device sends this image as digital data to the server.
[0536] Step 2:
[0537] The server processes the received image data using an optical character recognition (OCR) processing unit. The input is the transmitted image data, and the output is the extracted text information. The OpenCV library and other tools are used to extract the text information from the image.
[0538] Step 3:
[0539] The server analyzes the extracted text information using a problem analysis device. The input is text information, and the output is the identified problem and its classification result. The problem is identified by referring to a database and matching it with similar known problems.
[0540] Step 4:
[0541] The user's facial expressions and voice data are transmitted from the terminal to the server. The input is the user's facial expressions and voice data, and the output is a state ready for emotion analysis. The data is converted to an appropriate format for use by the emotion analysis device.
[0542] Step 5:
[0543] The server uses an emotion analysis device to identify the user's emotional state. Input is facial expressions and voice data, and output is the user's emotional state. Azure Cognitive Services is used to analyze the user's stress and concentration levels.
[0544] Step 6:
[0545] The server adjusts the learning content using an adaptive learning device based on the emotion analysis results. The input is the emotional state and identified task, and the output is the adjusted learning content. The difficulty and amount of the learning content are dynamically adjusted to match the user's state.
[0546] Step 7:
[0547] The server sends the tailored learning content to the device. The input is the learning material, and the output is the educational content displayed to the user. The device uses a display device to visually present the content in order to provide a learning experience optimized for the user.
[0548] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0549] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0550] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0551] [Fourth Embodiment]
[0552] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0553] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0554] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0555] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0556] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0557] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0558] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0559] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0560] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0561] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0562] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0563] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0564] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0565] The learning support system of the present invention operates by allowing users to capture exam questions and input answer information using a mobile device such as a smartphone or tablet. The user photographs the exam questions with a camera and indicates on the interface whether the answer is correct or incorrect. The captured image data and answer information are uploaded from the device to a server.
[0566] The server first performs optical character recognition (OCR) on the received image to extract text data. This process digitizes text information such as the question and answer choices. Next, the server analyzes the extracted text data, compares it with the registered learning database to identify and classify the question. Based on the identified question category and difficulty level, the server evaluates the user's answer results and analyzes areas that need to be studied.
[0567] The analysis results are sent from the server to the user's terminal. The terminal visually displays the received analysis results to the user. Based on these results, the user can identify their strengths and areas of learning that need strengthening, and efficiently develop a learning plan.
[0568] For example, if a user takes a picture of a differential equation problem from high school mathematics and records an incorrect answer, the server categorizes the problem as a differential equation. It then provides the user with advice indicating that they need to strengthen their understanding of differential equations, from fundamentals to applications. In this way, the present invention provides efficient learning support tailored to the individual learning needs of each user.
[0569] The following describes the processing flow.
[0570] Step 1:
[0571] The user takes a picture of the exam question using their smartphone camera. The user then reviews the image through an interface and selects whether their answer to the question is correct or incorrect.
[0572] Step 2:
[0573] The device packages the captured image and the user's entered answer information to send to the server, and uploads the data to the server via the network.
[0574] Step 3:
[0575] The server receives the uploaded image and uses OCR (Optical Character Recognition) technology to extract the text within the image. The extracted text includes the question and answer choices.
[0576] Step 4:
[0577] The server analyzes the text data obtained by OCR to identify and classify the problems. This includes matching the data against existing databases to determine the problem category and difficulty level.
[0578] Step 5:
[0579] The server considers the user's answer results (correct / incorrect) and analyzes the problem information to determine the optimal learning areas for the user. In particular, it identifies the user's weak areas and determines areas that need focused reinforcement.
[0580] Step 6:
[0581] The server compiles the analysis results, creates a report in a format that is easy for the user to understand, and sends that data to the terminal.
[0582] Step 7:
[0583] The device receives analysis results sent from the server and displays them visually for the user. The user can then use this information to adjust their learning plan.
[0584] (Example 1)
[0585] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0586] There is a need for a system that can more efficiently identify the learning challenges each user has and provide optimal learning support tailored to their individual learning needs. However, conventional learning support systems have problems such as insufficient collection and analysis of user learning data, resulting in inadequate individual learning suggestions.
[0587] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0588] In this invention, the server includes means for capturing information using an information acquisition device, optical recognition means for recognizing data from the captured information, and information analysis means for identifying and classifying information based on the recognized data. This enables detailed analysis and optimized learning support tailored to individual learning situations.
[0589] An "information acquisition device" is a device used by users to capture data and information, and generally includes cameras and scanners.
[0590] An "optical recognition means" is a method for recognizing characters and data from captured information and extracting them as digital information, and optical character recognition (OCR) technology is commonly used.
[0591] "Information analysis means" refers to methods for identifying and classifying information based on recognized data and performing analysis related to user learning.
[0592] An "analysis presentation tool" is a means of providing learners with visual or other feedback based on analyzed information, thereby supporting the optimization of their learning.
[0593] A "central processing unit" refers to a computing environment or server that operates in the background to receive and process data and generate analysis results.
[0594] The learning support system of the present invention operates by allowing the user to capture exam questions using a portable device, such as a smartphone or tablet, and input the answer information. The user uses the camera function of the portable device to take an image of the exam question. Next, the user inputs the answer through an interface within the device's application that specifies whether the answer is correct or incorrect.
[0595] The terminal transmits the captured image data and answer information to the server. The server uses optical character recognition software (e.g., Tesseract) to extract character data from the received image data. This enables the conversion of paper-based exam questions into digital data.
[0596] The server further analyzes the extracted text data, identifies similar information from the registered database, and classifies the problem. This analysis may utilize machine learning algorithms and database lookup techniques. After the analysis, the server evaluates the problem category and difficulty level, and analyzes which learning areas the user should strengthen. The final analysis results are sent to the user's terminal and provided as visual feedback.
[0597] For example, if a user takes a picture of a differential equation problem from high school mathematics and records their incorrect answer, the server analyzes the data and categorizes the problem as "differential equations." Based on the information obtained from the analysis, the server can provide the user with advice such as, "You need to strengthen your knowledge of differential equations, from basic to applied levels." This information helps the user to create an effective study plan.
[0598] An example of a prompt message would be, "Please suggest educational materials to deepen our understanding of the fundamentals of differential equations," which allows for further information gathering using a generative AI model.
[0599] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0600] Step 1:
[0601] The user uses the device to take a picture of the test questions with the camera. The image data is saved to the device as input. The user then inputs whether the answer is correct or incorrect on the device's interface. This information is associated with the captured image data, and the device prepares to send it to the server.
[0602] Step 2:
[0603] The device uploads the captured image data and associated answer information to the server. The image data and answer are transferred as input, and the server receives and stores this information. Secure protocols (such as HTTPS) are used for communication to ensure data security.
[0604] Step 3:
[0605] The server performs optical character recognition (OCR) processing on the received image data. The input is image data sent from the terminal, and the output is string data. Specifically, it uses OCR software (e.g., Tesseract) to convert the character information in the image into text.
[0606] Step 4:
[0607] The server analyzes the character data generated by OCR and compares it with a training database. The input is string data from OCR, and the output identifies the problem category and difficulty level. In this process, the server uses a text analysis algorithm to classify the problem.
[0608] Step 5:
[0609] The server evaluates the user's answers based on the analysis results. The input consists of categorized problem data and the user's answers, and the output proposes individual learning reinforcement areas. The server performs a statistical evaluation to analyze which areas the user needs to improve their knowledge in.
[0610] Step 6:
[0611] The server sends the analysis results to the user's terminal. The input is the analyzed data, and the output is specific learning advice for the user. The server structures this information and transfers it to the terminal.
[0612] Step 7:
[0613] The terminal visually presents the analysis results received from the server to the user. The input is the analysis data sent from the server, and the output is an information display that is intuitively understandable to the user. Specifically, it visualizes learning advice using graphs and charts.
[0614] (Application Example 1)
[0615] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0616] Early detection of abnormalities in equipment and components within a factory, and efficient identification of their causes, are crucial challenges in maintaining manufacturing lines. Conventional methods often involve delays in detecting abnormalities, leading to delayed responses. These challenges need to be addressed.
[0617] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0618] In this invention, the server includes an image capture means, a recognition means for recognizing information from the captured image, and an analysis means for identifying and classifying events based on the recognized information. This makes it possible to quickly detect abnormalities in equipment and parts within a factory, identify their causes, and prompt appropriate countermeasures.
[0619] "Image acquisition means" refers to the functions of devices or equipment used to capture images of a subject.
[0620] "Recognition means" refers to a function for extracting and identifying information from captured images.
[0621] "Analysis means" refers to a function for identifying and classifying events based on recognized information.
[0622] "Presentation means" refers to a function for providing the analyzed information to the user.
[0623] "Means for providing feedback through recognition and analysis means mounted on the device" refers to functions that provide rapid feedback using functions built into the device.
[0624] This system utilizes smart glasses and a cloud server. The smart glasses, with their built-in camera and interface, support on-site operations. Specifically, operators use the smart glasses to photograph abnormalities in machinery and equipment within the factory. This image data is transmitted to the cloud server via the internet. The server extracts text data from the images using OCR technology and further identifies and classifies abnormalities using AI analysis such as TensorFlow. The analysis results are fed back to the smart glasses via the cloud, visually presenting the cause of the abnormality and countermeasures to the user. This enables rapid problem solving within the factory.
[0625] For example, if an operator detects an anomaly in a conveyor belt on a manufacturing line, they can use smart glasses to photograph the area. This image is sent to a cloud server, where AI analysis diagnoses that dirt has accumulated on the belt's sensor. The feedback includes specific instructions such as, "Please clean the conveyor belt sensor."
[0626] An example of a prompt message would be: "This image shows a part of a factory's conveyor line. If there is an abnormality, please tell me the cause and how to address it." Using this prompt message, the AI model derives appropriate analysis results.
[0627] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0628] Step 1:
[0629] The user uses smart glasses to photograph abnormal areas on machinery and equipment within the factory. The input is image data captured by the smart glasses' camera. This image data is temporarily stored in the terminal's memory for the next processing step.
[0630] Step 2:
[0631] The device uploads captured image data to a cloud server. The input is image data stored on the device, and the output is data sent to the server via the internet. This involves secure communication using security protocols such as SSL.
[0632] Step 3:
[0633] The server extracts text data from the received image data using OCR technology. The input is image data, and the output is extracted text information. The OCR library (e.g., Tesseract) recognizes the text within the image.
[0634] Step 4:
[0635] The server uses OCR to extract text information, which is then analyzed using AI. The input is the text information obtained from OCR, and the output is the result of identifying and classifying anomalies. Using generative AI models such as TensorFlow, the text information is analyzed to identify the anomalous parts and their causes.
[0636] Step 5:
[0637] The server feeds the analysis results back to the smart glasses terminal. The input is the result of the AI analysis, and the output is the feedback information sent to the user terminal. The data is transmitted over the network and displayed on the user interface.
[0638] Step 6:
[0639] The user receives feedback on anomalies through smart glasses and takes appropriate action. Based on the outputted information, the user implements specific countermeasures. For example, instructions such as "Please clean the conveyor belt sensor" are visually displayed.
[0640] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0641] The learning support system of this invention combines three main functions—image recognition, answer input, and emotion analysis—to maximize the user's learning efficiency. Users can use mobile devices such as smartphones and tablets to take pictures of exam questions and input their answers. In addition, the emotion engine analyzes the user's emotional state through the camera and microphone.
[0642] The terminal uploads images of the test questions and the user's answers to the server. At the same time, the user's facial expressions and voice data, captured by the device's camera and microphone, are also transmitted. The server first performs OCR on the images to extract text information. Next, a question analysis tool analyzes this text, comparing it to an existing database to identify and classify the questions.
[0643] Subsequently, the server uses an emotion engine to analyze the user's facial expressions and voice data to identify their emotional state. This emotion analysis is used to evaluate the user's stress level, concentration level, and other factors, and helps to adjust the learning presentation methods. Specifically, the difficulty and amount of learning content presented are adjusted based on the emotional state to optimize the learning experience and prevent the user from becoming overloaded.
[0644] For example, if the emotion engine detects signs of increased stress in the user when they select an incorrect answer to a difficult question, the server will adjust the difficulty level of the next learning question or display an encouraging message. In this way, the present invention enhances the quality of learning and provides personalized and appropriate learning support to the user.
[0645] The following describes the processing flow.
[0646] Step 1:
[0647] Users take photos of exam questions with their smartphone cameras and select whether their answers are correct or incorrect using an on-screen interface. The camera and microphone also capture the user's facial expressions and voice.
[0648] Step 2:
[0649] The device structures the data—including images, answer information, and acquired facial expressions and audio data—and uploads it to the server via the network in order to send it to the server.
[0650] Step 3:
[0651] The server receives the image data, first applying OCR technology to recognize the text within the image and extracting the character data of the question and answer choices.
[0652] Step 4:
[0653] The server performs problem analysis based on extracted text data, and identifies and classifies the problem category and difficulty level by referring to the platform's database.
[0654] Step 5:
[0655] The server uses an emotion engine to analyze received facial expression and voice data and evaluate the user's emotional state. This includes generating indicators such as stress levels and concentration levels.
[0656] Step 6:
[0657] The server combines the user's answer results with the sentiment analysis results and adjusts the learning content presented using a learning presentation method. Specifically, it optimizes the difficulty level and adds motivational messages.
[0658] Step 7:
[0659] The server compiles the adjusted learning suggestions into a report and sends it to the user's terminal.
[0660] Step 8:
[0661] The device receives analysis results from the server and presents them visually to the user. The user can then use this information to improve their learning plan.
[0662] (Example 2)
[0663] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0664] Conventional learning support systems fail to adjust learning content according to the user's emotional state, and in particular, they are insufficiently optimized for the learning experience in response to stress levels and concentration levels. Furthermore, they lack real-time feedback when users solve problems, making it difficult to improve learning efficiency.
[0665] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0666] In this invention, the server includes an image acquisition means, an optical character recognition means, an analysis means, an emotion analysis means, and an adjustment means. This makes it possible to extract and analyze character information from the user's image data, further analyze the user's emotional state, and dynamically adjust the learning content based on that.
[0667] "Image acquisition means" refers to a function that allows users to take pictures of learning materials using a mobile device.
[0668] "Optical character recognition means" refers to a technology that converts character information from image data into digital text.
[0669] "Analysis means" refers to a function that analyzes text information acquired by optical character recognition means, identifies the information, and classifies it.
[0670] "Presentation method" refers to a function that provides users with appropriate learning content based on analysis results and sentiment analysis results.
[0671] "Emotional analysis methods" refer to the process of analyzing a user's facial expressions and voice data to determine their emotional state.
[0672] The "adjustment mechanism" is a function that dynamically and appropriately changes the difficulty level and amount of learning content presented based on the emotion analysis results.
[0673] This learning support system operates through the collaboration of users, devices, and a server. Users can use devices such as smartphones and tablets to capture images of exam questions and input answer data during their studies. The device then sends the captured image data and answer data to the server.
[0674] The server extracts character information as digital text from received image data using high-precision optical character recognition (OCR) software. This technology can utilize existing OCR engines, such as Tesseract. The extracted text is analyzed by analysis software, which compares it against existing databases to identify and classify problems. Database management systems and natural language processing libraries are suitable for this analysis.
[0675] Next, the server supplies the user's facial expressions and voice data acquired through the camera and microphone to the emotion analysis engine to identify the user's emotional state. This analysis may utilize machine learning models, such as those using "OpenCV" or "Librosa."
[0676] The server adjusts the learning presentation methods based on the sentiment analysis results. Specifically, it changes the difficulty level and amount of learning content presented, and displays encouraging messages to the user, enabling them to continue learning efficiently. This optimizes the learning experience for each individual.
[0677] For example, if a user answers a difficult question incorrectly and the server detects an increase in stress through emotion analysis, it can maintain the user's motivation by making the next learning question easier or by providing an encouraging message.
[0678] An example of a prompt generated by a generative AI model is, "Please describe a system that adjusts learning content by considering the user's emotional state in order to maximize the user's learning efficiency."
[0679] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0680] Step 1:
[0681] Users take photos of exam questions using their mobile devices and input their answers. During this process, the device's camera and microphone simultaneously capture the user's facial expressions and voice. The input consists of the captured image, answer data, facial expression data, and voice data. This data is stored on the device and then prepared for transmission to the server.
[0682] Step 2:
[0683] The terminal sends image data, answer data, facial expression data, and audio data obtained from the user to the server. The input is the aforementioned dataset. For secure data transfer, it is desirable to use HTTPS as the protocol for transmission. As output, a receipt confirmation is displayed on the terminal.
[0684] Step 3:
[0685] The server processes the received image data into optical character recognition (OCR) software. The input is image data of the exam questions. The OCR process extracts the character information as digital text, which is then passed to the question analysis software. The output is the recognized character information.
[0686] Step 4:
[0687] The server analyzes text information, compares it with a database, and identifies and classifies the problems. The input is text information obtained by OCR. In this process, the type and difficulty level of the problems are identified, and they are classified as learning content. The output is the identified problem type and classification information.
[0688] Step 5:
[0689] The server passes the received facial expression and voice data to the emotion analysis engine. The input is biometric data. The emotion analysis engine processes the data to identify the user's emotional state, particularly stress levels and concentration levels. The output is information about the user's emotional state.
[0690] Step 6:
[0691] The server adjusts the presentation of learning content based on analyzed problem and emotional state information. Inputs are identified problem classifications and emotional states. Based on this data, the server adjusts the difficulty and amount of the next learning content presented, or generates encouraging messages. Outputs are the adjusted learning plan and feedback.
[0692] Step 7:
[0693] The server sends a customized learning plan and feedback to the user's device. The input consists of generated learning content and messages. Based on the received information, the device provides a learning experience optimized for the user. The output is the learning content and feedback displayed on the user's screen.
[0694] (Application Example 2)
[0695] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0696] In today's world, there is a demand for flexible learning support tailored to each user's individual learning pace and emotional state, but conventional systems have struggled to achieve this in real time and effectively. In particular, optimizing learning content while considering the user's emotional state is a challenge.
[0697] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0698] In this invention, the server includes an image acquisition device, a character recognition processing device, and a task analysis device. This enables dynamic adjustment of the learning content according to the user's answer information and emotional state.
[0699] An "image acquisition device" is a device that captures visual information in a specific environment, and includes cameras and scanners incorporated into digital devices.
[0700] A "character recognition processing device" is a processing device that uses optical methods to read character information from acquired image data and digitizes it.
[0701] A "task analysis device" is a device that has the function of identifying and classifying specific tasks or problems based on information obtained from a character recognition processing device.
[0702] A "learning display device" is a device that presents users with tailored learning content and has functions to promote an efficient learning experience.
[0703] An "emotion analysis device" is a device that analyzes a user's facial expressions and voice data to identify their emotional state.
[0704] An "adaptive learning device" is a device that uses the results of an emotion analysis device to dynamically adjust the difficulty level and amount of learning content to be optimal for the user.
[0705] To realize this invention, it is necessary to combine multiple hardware and software components. The system consists of a user's mobile terminal, a server, and various analysis programs.
[0706] The user's mobile device is equipped with a camera as an image acquisition device, which can be used to photograph learning materials and problems. This image data is transmitted to a server via a communication means. On the server, image recognition software is executed for optical character recognition, such as OpenCV. This software processes the extraction of character information from the image.
[0707] After character recognition is complete, the next component designed is a problem analysis unit. Here, the extracted text data is analyzed to identify and classify the problem. A database is used in this process to match the problem with similar issues. Next, an emotion analysis unit analyzes the received user's facial expressions and voice data to determine their emotional state. Emotion analysis APIs such as Azure Cognitive Services are utilized for this purpose.
[0708] The results of the emotion analysis are sent to the adaptive learning device and used to adjust the learning content. For example, if the user is experiencing stress, the difficulty level of the task is adjusted. The content adjusted by the learning display device is then sent to the user's terminal, providing the user with an optimal learning experience.
[0709] As a concrete example, suppose a user is learning a mathematical formula. When they take a picture of the material with their camera, a simple question related to that formula is presented. Sentiment analysis is performed on the user's answer, and in some cases, an encouraging message is displayed.
[0710] An example of a prompt for a generative AI model is: "Consider how to adjust the learning content to match situations where the user feels burdened. Propose a method for analyzing the user's emotional state and appropriately adjusting the difficulty level."
[0711] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0712] Step 1:
[0713] The user takes a picture of the learning material using the camera on their mobile device. The input is an image of the learning material, and the output is image data. The device sends this image as digital data to the server.
[0714] Step 2:
[0715] The server processes the received image data using an optical character recognition (OCR) processing unit. The input is the transmitted image data, and the output is the extracted text information. The OpenCV library and other tools are used to extract the text information from the image.
[0716] Step 3:
[0717] The server analyzes the extracted text information using a problem analysis device. The input is text information, and the output is the identified problem and its classification result. The problem is identified by referring to a database and matching it with similar known problems.
[0718] Step 4:
[0719] The user's facial expressions and voice data are transmitted from the terminal to the server. The input is the user's facial expressions and voice data, and the output is a state ready for emotion analysis. The data is converted to an appropriate format for use by the emotion analysis device.
[0720] Step 5:
[0721] The server uses an emotion analysis device to identify the user's emotional state. Input is facial expressions and voice data, and output is the user's emotional state. Azure Cognitive Services is used to analyze the user's stress and concentration levels.
[0722] Step 6:
[0723] The server adjusts the learning content using an adaptive learning device based on the emotion analysis results. The input is the emotional state and identified task, and the output is the adjusted learning content. The difficulty and amount of the learning content are dynamically adjusted to match the user's state.
[0724] Step 7:
[0725] The server sends the tailored learning content to the device. The input is the learning material, and the output is the educational content displayed to the user. The device uses a display device to visually present the content in order to provide a learning experience optimized for the user.
[0726] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0727] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0728] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0729] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0730] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0731] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0732] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0733] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0734] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0735] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0736] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0737] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0738] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0739] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0740] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0741] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0742] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0743] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0744] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0745] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0746] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0747] The following is further disclosed regarding the embodiments described above.
[0748] (Claim 1)
[0749] Image capture method,
[0750] An optical character recognition means for recognizing characters from captured images,
[0751] A problem analysis method that identifies and classifies problems based on recognized characters,
[0752] A learning presentation method that presents learning areas based on the user's answer results,
[0753] A learning support system that includes this.
[0754] (Claim 2)
[0755] The learning support system according to claim 1, further comprising an answer input means for inputting the answer result selected by the user.
[0756] (Claim 3)
[0757] The learning support system according to claim 1, characterized in that the optical character recognition means is executed on a server.
[0758] "Example 1"
[0759] (Claim 1)
[0760] Means for capturing information using an information acquisition device,
[0761] An optical recognition means for recognizing data from captured information,
[0762] Information analysis means for identifying and classifying information based on recognized data,
[0763] An analytical presentation method that evaluates and presents the learning domain based on the analyzed information,
[0764] A system characterized by data processing being performed by a central processing unit.
[0765] (Claim 2)
[0766] The system according to claim 1, further comprising an information input means for inputting answer data selected by the user.
[0767] (Claim 3)
[0768] The system according to claim 1, comprising a presentation device that provides visual feedback to the user based on the analyzed information.
[0769] "Application Example 1"
[0770] (Claim 1)
[0771] Image capture method,
[0772] A recognition means for recognizing information from captured images,
[0773] An analytical means for identifying and classifying events based on recognized information,
[0774] A means of presenting information based on the user's results,
[0775] A means for providing feedback by recognition means and analysis means mounted on the device,
[0776] A system that includes this.
[0777] (Claim 2)
[0778] The system according to claim 1, further comprising an input means for inputting a result selected by the user.
[0779] (Claim 3)
[0780] The system according to claim 1, characterized in that the analysis means is executed on a server.
[0781] "Example 2 of combining an emotion engine"
[0782] (Claim 1)
[0783] Image acquisition method,
[0784] An optical character recognition means for recognizing characters from acquired images,
[0785] An analysis method that identifies and classifies information based on recognized characters,
[0786] A presentation method that presents learning content based on user response data,
[0787] An emotion analysis means that analyzes the user's facial expressions and voice data to determine their emotional state,
[0788] An adjustment mechanism that adjusts the difficulty level and amount of learning content based on the identified emotional state,
[0789] A system that includes this.
[0790] (Claim 2)
[0791] The system according to claim 1, further comprising an input means for inputting answer data selected by the user.
[0792] (Claim 3)
[0793] The system according to claim 1, characterized in that the optical character recognition means is performed on a data processing device.
[0794] "Application example 2 when combining with an emotional engine"
[0795] (Claim 1)
[0796] Image acquisition device,
[0797] A character recognition processing device that extracts information from acquired images,
[0798] A problem analysis device that identifies and classifies problems based on extracted information,
[0799] A learning display device that adjusts learning content based on the user's answer information,
[0800] An emotion analysis device that analyzes the emotional state of the user,
[0801] An adaptive learning device that dynamically adjusts the difficulty level and amount of learning content based on the results of emotion analysis,
[0802] A system that includes this.
[0803] (Claim 2)
[0804] The system according to claim 1, further comprising an answer receiving device that receives answer information selected by the user.
[0805] (Claim 3)
[0806] The system according to claim 1, characterized in that the character recognition processing device is executed on a communication device. [Explanation of Symbols]
[0807] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. Image capture method, An optical character recognition means for recognizing characters from captured images, A problem analysis method that identifies and classifies problems based on recognized characters, A learning presentation method that presents learning areas based on the user's answer results, A learning support system that includes this.
2. The learning support system according to claim 1, further comprising an answer input means for inputting the answer result selected by the user.
3. The learning support system according to claim 1, characterized in that the optical character recognition means is executed on a server.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A