system
A system analyzes incorrectly answered problems to identify student weaknesses and provides personalized learning plans, improving educational efficiency by addressing the limitations of conventional methods.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-11-13
- Publication Date
- 2026-05-25
AI Technical Summary
Conventional educational methods fail to identify individual student weaknesses effectively and provide tailored learning plans, leading to inefficient learning experiences, especially among junior and senior high school students and examinees.
A system that analyzes images of incorrectly answered problems to extract text data, identifies weaknesses, generates similar problems, provides detailed explanations, and records learning progress to create personalized learning plans.
Enhances learning efficiency by allowing students to focus on their weaknesses and receive customized educational support, resulting in more effective learning outcomes.
Smart Images

Figure 2026085775000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the modern educational environment, especially among junior and senior high school students and examinees, there is a problem that incorrect questions are often left as they are and effective review is not carried out. In addition, it is difficult to identify the weak fields according to individual students and provide appropriate learning plans based on them, and there is a problem that the conventional educational methods cannot flexibly respond to the individual needs of students.
Means for Solving the Problems
[0005] This invention provides a system that extracts text data from learning problems by analyzing images taken by users of learning problems they have answered incorrectly. Furthermore, it analyzes the user's weak areas based on this data and generates and provides similar problems individually. It then deepens the user's understanding by providing detailed explanations for the generated similar problems. The system also includes means for recording learning progress, creating a new learning plan based on the user's learning history, and providing an individually optimized learning environment.
[0006] A "user" refers to an individual or legal entity that uses the system to learn and submits photos of incorrect answers.
[0007] "Image acquisition means" refers to a device or technology for recording learning problems that a user has answered incorrectly as digital data and transmitting that data to a system.
[0008] "Image analysis means" refers to a device or technology for recognizing characters and mathematical formulas from acquired image data and extracting them as text data.
[0009] "Analysis means" refers to a device or technology that analyzes the user's learning history and trends based on extracted text data and identifies areas of weakness.
[0010] "Problem generation means" refers to a device or technology that automatically creates and provides similar problems tailored to each user based on their areas of weakness.
[0011] "Explanatory provision means" refers to a device or technology for creating explanatory information related to similar problems that have been generated and presenting it to the user visually or audibly.
[0012] "Feedback provision means" refers to a device or technology for recording a user's learning progress, generating a new learning plan based on that data, and providing it to the user. [Brief explanation of the drawing]
[0013] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]
[0014] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0015] First, the terms used in the following description will be explained.
[0016] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0017] In the following embodiments, the signed RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0018] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0019] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0021] [First Embodiment]
[0022] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0023] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0024] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0025] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0026] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0028] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0029] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0030] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0031] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0032] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0033] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0034] This system consists of a user terminal and a server that performs analysis and data processing. First, the user takes photos of problems they answered incorrectly during their learning process using the terminal's camera function. This terminal has the capability to upload the image data captured by the user directly to the server.
[0035] The server uses OCR (Optical Character Recognition) technology to extract text data from received images for analysis. This data is stored in the server's data storage and used for subsequent analysis. Through this analysis, the server considers information about problems the user has previously answered incorrectly and how they were categorized, identifying areas where the user has particular weaknesses. This provides a foundation for generating appropriate similar problems tailored to the user's learning progress.
[0036] Next, the server generates similar problems tailored to the user's areas of weakness using a problem generation system. These problems are adjusted in difficulty and selected from an extensive problem database. Furthermore, the explanation provision system creates and provides explanations corresponding to the generated problems in text or video format. These explanations are designed to explain complex concepts in an easy-to-understand manner.
[0037] Users can receive these problems and explanations via their devices and review them. Furthermore, the user's problem-solving process and results are transmitted to the server in real time. The server records this information through a feedback system and updates the user's learning history. Based on these results, a new learning plan is generated to assist with future learning.
[0038] By implementing this entire system, users can efficiently learn to overcome their weak areas, resulting in more effective learning outcomes.
[0039] The following describes the processing flow.
[0040] Step 1:
[0041] Users take photos of problems they answered incorrectly during their studies using the camera on their smartphone or tablet. It is important that the problems are clearly and distinctly visible in these photos.
[0042] Step 2:
[0043] The device saves the captured image data to its own storage and uploads it to the server with user identification information attached. The device also checks the communication status to ensure that the data is transmitted correctly.
[0044] Step 3:
[0045] The server receives the uploaded image data and uses OCR technology to extract text data from the image. The server then uses the extracted characters and mathematical formulas to identify the problem statement and structures the information as text data.
[0046] Step 4:
[0047] The server analyzes the text data and compares it with the user's past learning history to identify areas or topics where the user struggles. In this process, it utilizes multiple databases to analyze and classify the user's answer patterns.
[0048] Step 5:
[0049] The server generates similar problems using a problem generation mechanism based on identified areas of weakness. This includes a process of selecting the difficulty level and format of the problems, and the appropriate problems are selected dynamically.
[0050] Step 6:
[0051] The server creates explanations for problems and prepares the content in text or video format. The explanations are provided in a way that includes relevant information and examples to aid understanding.
[0052] Step 7:
[0053] The terminal receives similar problems and explanations sent from the server and provides an interface for visualization to the user. The user can then view this information for learning and review.
[0054] Step 8:
[0055] Users input the results of solving the provided problems and their learning progress into their device and send this information back to the server. This input includes data on whether their answers were correct or incorrect.
[0056] Step 9:
[0057] The server evaluates the user's learning progress based on their input data and creates the next learning plan. Based on these results, it prepares new similar problems and review plans to be used in the next learning session.
[0058] (Example 1)
[0059] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0060] Conventional learning support systems have struggled to efficiently identify individual users' areas of weakness and generate and provide appropriate learning problems based on those areas. Furthermore, the lack of immediate feedback on the provided learning problems and the subsequent updates to learning plans led to decreased user learning efficiency.
[0061] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0062] In this invention, the server includes information acquisition means, information analysis means, and evaluation means. This makes it possible to photograph learning tasks that the user has answered incorrectly, analyze the information to extract textual information, and further identify the user's areas of weakness based on the extracted textual information.
[0063] "Information acquisition means" refers to a device or function that allows a user to take a picture of a learning task and transmit that data to a server.
[0064] "Information analysis means" refers to a device or function that analyzes received data and provides technology to accurately extract textual information from image information.
[0065] "Evaluation means" refers to a device or function that analyzes and identifies areas of weakness for the user based on extracted textual information.
[0066] "Problem generation means" refers to a device or function that creates and provides appropriate similar problems according to the user's areas of weakness.
[0067] "Explanation provision means" refers to a device or function that prepares a detailed explanation for a generated problem and presents it to the user.
[0068] "Evaluation provision means" refers to a device or function that records the user's learning progress and creates a new learning plan based on that progress.
[0069] This invention provides an educational support system that enables users to efficiently overcome learning tasks they have made mistakes on. The system consists of a user terminal and a server that performs data analysis and processing.
[0070] Users use their device's camera function to take pictures of problems they answered incorrectly during their studies. This device is equipped with a means to upload the image data directly to a server. For example, a smartphone or tablet can be used as the device.
[0071] The server extracts text information from captured images using OCR (Optical Character Recognition) technology as an information analysis method. This process utilizes technologies such as "Tesseract," which is widely used as OCR software.
[0072] Furthermore, the server analyzes the extracted character information using evaluation tools to identify the user's weak areas. By analyzing past learning data and accuracy rates, the user's areas of difficulty are clearly defined.
[0073] Using a task generation method, the server creates similar tasks based on the user's areas of weakness using a generation AI model. Technologies such as the "OpenAI® API" are utilized in this generation AI model. This model has the functionality to adjust the task difficulty to an appropriate level, taking into account the user's learning history.
[0074] The generated tasks are accompanied by detailed explanations provided through explanatory tools. These explanations are delivered to users in either text or video format. Tools such as "Explain Everything" are used for video explanations, allowing users to intuitively deepen their understanding.
[0075] Finally, using feedback mechanisms, the server records the user's learning progress and automatically generates a new learning plan. This plan optimizes subsequent learning content based on the user's proficiency level.
[0076] As a concrete example, consider a scenario where a user struggles with a fraction problem in mathematics. In this case, the user takes a picture of the problem and uploads it to the server. The server analyzes the image, identifies the user's lack of understanding related to fractions, and then generates similar practice problems and prepares detailed explanations. An example of a prompt to be input to the generation AI model would be: "Generate similar problems and explanations to help the user solve the fraction problem they struggle with. Please provide concrete examples and consider ways to explain complex concepts simply."
[0077] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0078] Step 1:
[0079] During the learning process, users take photos of problems they answered incorrectly using their device's camera. A dedicated application is installed on the device, allowing it to directly send image data to a server. The input is the image data of the problems the user has photographed. The output is the uploading of that data to the server. By selecting the appropriate problems and photographing them with the camera, users provide accurate information for the system to analyze.
[0080] Step 2:
[0081] The server extracts text information from received image data using OCR technology. The input is the image data uploaded to the server. The output is the extracted text data. The server uses OCR software such as "Tesseract" to read the text information in the image and convert it into parseable text. This process clarifies the specific content of the questions the user answered incorrectly.
[0082] Step 3:
[0083] The server analyzes the extracted text data to identify the user's areas of weakness. The input is the extracted text data. The output is a list of the user's areas of weakness. Past learning data and accuracy rates are used in the analysis, and statistical methods are employed to identify the user's weak areas. This enables personalized educational support.
[0084] Step 4:
[0085] The server generates similar problems based on the identified areas of weakness. A generative AI model is used for this process. The input is a list of the user's areas of weakness. The output is the generated similar problems. AI technologies such as the "OpenAI API" are utilized to create appropriate tasks that deepen the user's understanding. The generated problems are tailored to the user's specific learning needs.
[0086] Step 5:
[0087] The server prepares explanations for the generated problems. The input is a similar problem that has been generated. The output is a detailed explanation of the problem, which includes text and video explanations. Tools such as "Explain Everything" are used to create easy-to-understand explanations, helping users to understand the problems more deeply.
[0088] Step 6:
[0089] The user reviews the material based on the questions and explanations received on their device. During this process, the user's answers are sent to the server in real time. Input consists of the user's answers and their results. Output is the server-side update of the learning progress data. The device records the user's answer history and provides it to the server as feedback.
[0090] Step 7:
[0091] The server updates the user's learning plan based on the feedback provided. The input is the updated learning progress data. The output is the new learning plan. The plan generated by the server is used in the user's next learning session. This process optimizes continuous learning.
[0092] (Application Example 1)
[0093] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0094] In today's information society, personalized learning support tailored to each user's learning characteristics is required to respond quickly and accurately to individual learning needs. However, existing learning systems have the challenge of not being able to accurately identify a user's weak areas and provide effective learning content. To solve this problem, an advanced system is needed that analyzes the user's learning history in real time and provides personalized feedback and learning plans.
[0095] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0096] In this invention, the server includes data acquisition means for taking and uploading photos of learning problems that the user has answered incorrectly; data analysis means for analyzing the uploaded images and extracting text data; information analysis means for identifying the user's weak areas from the extracted text data; and data processing means for generating easy-to-learn visual video content based on the analysis results of the user's weaknesses and distributing it to a data device. This enables the provision of learning content tailored to each individual user and facilitates efficient understanding.
[0097] A "data acquisition means" is a device equipped with the function of taking a picture of a learning problem that the user answered incorrectly and uploading that data to a server.
[0098] A "data analysis means" is a device that has the technology to analyze uploaded images and extract text data from those images.
[0099] An "information analysis tool" is a device that performs data processing based on extracted text data to identify areas where the user struggles.
[0100] An "information generation means" is a device that has the function of creating and providing similar problems based on the user's areas of weakness.
[0101] An "information provision device" is a device equipped with the function of preparing and presenting explanations for similar problems that have been generated to the user.
[0102] A "results-providing device" is a device that records the user's learning progress and creates a new learning program based on that data.
[0103] A "data processing means" is a device that generates visual video content based on the analysis results of the user's weaknesses and distributes it to a data device.
[0104] Embodiments of the present invention are systems for improving user learning efficiency. This system consists of a terminal that the user normally uses and a server that is responsible for analysis and data processing.
[0105] The user first takes a picture of the learning problem they answered incorrectly using the device's camera function. This device has a built-in data acquisition mechanism that allows the user to upload the captured image directly to the server. The server then uses OCR (Optical Character Recognition) technology to analyze the uploaded image and extract the text data. Specifically, an OCR tool such as Tesseract is used.
[0106] The server identifies the user's areas of difficulty based on this text data and generates personalized information using information analysis tools. In this process, it selects appropriate similar problems for the user and adjusts their difficulty level. It also prepares explanations based on user-specific information and presents them to the user through information delivery tools.
[0107] Furthermore, the server records learning progress and presents the user with a new learning program using a results delivery system. The user's weaknesses are analyzed, and based on the results, visually easy-to-understand video content is generated. The generated content is distributed to a data device via a data processing system.
[0108] For example, a user who struggles with an English grammar problem will be shown an animated explanation of that grammatical point to help them solve the problem. Through this process, users can effectively engage in learning.
[0109] As a concrete example, here are some examples of prompt statements for a generative AI model.
[0110] "When solving the following grammar problems, identify common mistakes and generate related explanations."
[0111] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0112] Step 1:
[0113] The user takes a picture of the learning problem with their camera and saves it to their device.
[0114] The input is an image of the problem taken by the user. The user uploads this image to the server using a data acquisition method.
[0115] Step 2:
[0116] The device uploads the captured image to the server.
[0117] The input is captured image data, and the image data is sent to the server as output. This gives the server the foundation to receive the image data and begin processing it.
[0118] Step 3:
[0119] The server uses OCR technology to extract text data from the image.
[0120] The input is image data uploaded to the server, and the output is extracted text data. In this process, an OCR tool such as Tesseract is used to convert the text information in the image into digital text.
[0121] Step 4:
[0122] The server analyzes the extracted text data to identify the user's areas of weakness.
[0123] The input is extracted character data, and the output is information about the user's areas of difficulty. The server uses information analysis tools to analyze the patterns shown in the data and determine whether the user is repeatedly making mistakes in a particular area.
[0124] Step 5:
[0125] The server generates similar problems based on the user's areas of weakness.
[0126] The input is information on areas of weakness, and the output is a generated similar problem. The server uses information generation tools to select appropriate problems from a broad problem database and generates similar problems while adjusting the difficulty level.
[0127] Step 6:
[0128] The server prepares explanations for similar problems that have been generated and distributes them to users.
[0129] The input is a generated similar problem, and the output is related explanatory content. The server uses information delivery methods to create explanations in text or video format, and efforts are made to ensure that the explanations are easy for the user to understand.
[0130] Step 7:
[0131] The server records the user's learning progress and creates a new learning program.
[0132] The input is the user's problem-solving history and its analysis results, and the output is a learning program based on the next learning stage. The server maintains the learning record using a results-providing mechanism and presents the user with appropriate feedback and the next learning plan based on it.
[0133] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0134] This invention embodies a system that provides a personalized learning experience that takes into account the user's emotional state when they are learning incorrectly. The system comprises a terminal equipped with an emotion engine, a server that performs analysis processing, and a user interface.
[0135] First, the user takes a picture of the questions they answered incorrectly during the learning process using the device's camera. The device uploads these images to a server. Additionally, a built-in emotion engine monitors the user's voice and facial expressions in real time. From this data, the system detects the user's current emotional state.
[0136] The server analyzes the received image data using OCR technology and extracts the text data for the problem. Based on this data, it compares it with the user's past learning history to identify areas of difficulty. By also considering emotional data in identifying areas of difficulty, it becomes possible to estimate the user's stress level and motivation and present appropriate learning materials.
[0137] In generating similar problems, the server adjusts the difficulty level based on the user's emotional data. For example, if the server determines that the user is stressed, it can start with easier problems and gradually increase the difficulty. This reduces the burden of learning and allows for more effective review.
[0138] The explanations can be flexibly adapted to user sentiment data to support their understanding. For example, short text can be provided when the user is relaxed, while detailed video content can be offered when the user is highly focused.
[0139] Users progress through the provided similar problems and explanations, and upon completing a problem, the results are sent back to the server from their device. The server records these results as feedback and evaluates the user's progress. The newly generated learning plan is used for the next learning session, continuously providing a learning experience optimized for the user's individual needs.
[0140] For example, if a user makes a mistake on a math problem, the emotion engine detects the user's level of fatigue as soon as they upload the image they took. Based on this emotion information, the server incorporates relaxing problems and videos into the review plan. As a result, the user can learn efficiently while reducing stress.
[0141] The following describes the processing flow.
[0142] Step 1:
[0143] The user takes a picture of the question they answered incorrectly using their device's camera. The image should clearly show the question, the answer choices, and the answer the user selected.
[0144] Step 2:
[0145] The device saves captured images to its storage and uploads them to a server along with the user's identification information. Furthermore, it activates an emotion engine to analyze the user's facial expressions and voice in real time and collect emotion data.
[0146] Step 3:
[0147] The server analyzes the received image data using OCR technology and extracts text data from the image. At this time, the question, answer choices, and answer information are extracted as text.
[0148] Step 4:
[0149] The server analyzes the extracted text data, compares it with the user's learning history, and identifies areas of weakness. This analysis reflects the user's past performance and learning patterns.
[0150] Step 5:
[0151] The server generates similar problems while taking emotional data into consideration. If the user's emotional state is relaxed, it selects more difficult problems; if they are stressed, it prioritizes easier problems.
[0152] Step 6:
[0153] The server prepares explanations related to similar problems that have been generated. If the user's level of understanding is deemed high, a detailed explanation is provided; if their concentration is low, a concise explanation is provided.
[0154] Step 7:
[0155] The terminal receives similar problems and explanations sent from the server and presents them to the user. The user can solve the problems and learn by referring to the explanations.
[0156] Step 8:
[0157] Users input the results of solving new problems into their devices and send that data to the server. This input may include accuracy rates and solving speed.
[0158] Step 9:
[0159] The server combines user response data and sentiment data to evaluate learning progress and generate a next learning plan. This plan is tailored to the user's specific learning style and circumstances.
[0160] (Example 2)
[0161] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0162] Learners sometimes abandon problems they got wrong during the learning process. Furthermore, stress and decreased motivation due to monotonous learning hinder learning efficiency. While there is a need to provide learning experiences tailored to each learner's level of understanding and emotional state, current systems struggle to meet these requirements. Therefore, this invention aims to improve learning efficiency and sustainability by providing a flexible and effective learning process that takes the user's emotional state into consideration.
[0163] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0164] In this invention, the server includes data analysis means, emotion analysis means, and task generation means. This makes it possible to provide learners with personalized learning problems and explanations that take into account the user's areas of weakness and emotional state when they make mistakes on a problem.
[0165] "Data acquisition means" refers to a function that allows users to take photos of learning problems they answered incorrectly, collect the image data, and upload it to the server.
[0166] "Data analysis means" refers to a function that analyzes uploaded image data and extracts text information using optical character recognition.
[0167] "Emotional analysis means" refers to a function that analyzes data such as the user's voice tone and facial expressions to evaluate the user's emotional state in real time.
[0168] The "analysis method" is a function that identifies areas of difficulty for the user based on extracted text information and emotional state.
[0169] The "task generation method" is a function that generates similar problems based on identified areas of weakness and emotional states, and provides them to the user.
[0170] "Information provision means" refers to a function that prepares explanations of similar problems that have been generated and presents them in a format that is adjusted according to the user's emotional state.
[0171] A "progress management tool" is a function that records the user's learning progress and creates a new learning plan that will be useful for future learning.
[0172] The present invention is a system for providing a personalized learning experience that takes into account the user's emotional state when the user learns through incorrect problems. The system includes a terminal equipped with emotion analysis capabilities, a server that performs analysis processing, and a user interface.
[0173] First, the user uses the device's camera to take pictures of the questions they answered incorrectly during their studies. The device uploads this captured image data to the server. The device also has a built-in emotion engine that acquires voice and facial expression data to determine the user's emotional state in real time. This emotion data is also sent to the server.
[0174] Next, the server analyzes the received image data using OCR (Optical Character Recognition) technology to extract the text information in question. The OCR technology used here is software that enables accurate analysis of text data. The extracted text data is then compared with the user's past learning history to identify areas where the user struggles. Furthermore, emotional data is taken into consideration during this identification process, allowing the user's stress level and motivation to be inferred. This makes it possible to present the user with learning materials that are optimal for them.
[0175] Furthermore, the server adjusts the difficulty level of the problems based on the user's emotional data. Specifically, if the user is feeling stressed, it starts with easier problems to reduce the burden. Explanations for the generated problems are provided flexibly according to the user's emotional state to aid their understanding. For example, if the user is relaxed, a concise text-based explanation is provided, while if they are highly focused, detailed video content is offered.
[0176] As a concrete example, consider a scenario where a user makes a mistake on a math problem. When the user sends an image taken with their device to the server, the emotion engine detects the user's level of fatigue. Based on this emotional information, the server proposes a review plan that includes relaxing content. For example, it might include simple math problems or visually relaxing content.
[0177] An example of a prompt for a generative AI model is, "Please write a simple explanation of the following math problem, assuming the user is in a relaxed state." This prompt allows for the customization of the explanation provided to the user, thereby supporting improved learning efficiency.
[0178] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0179] Step 1:
[0180] The user takes a picture of the question they answered incorrectly using the camera. The captured image data becomes the input. The device temporarily stores this image data and prepares to upload it to the server. The specific actions here are activating the camera, capturing the image, and saving the data.
[0181] Step 2:
[0182] The device uploads the captured image data to the server. At this time, the image data is input and sent to the server. Once the server receives the image data, data analysis becomes possible. Specifically, this process involves data transfer over the network and confirmation of receipt.
[0183] Step 3:
[0184] The server analyzes the received image data using OCR technology. The input is image data, and the output is extracted text information. The server uses OCR software to extract character information from the image data and convert it into text data. Specifically, processes such as image preprocessing, character recognition, and text conversion are performed.
[0185] Step 4:
[0186] The server processes the user's voice and facial expression data and performs emotion analysis. Voice and facial expression data are inputs, and the output is the user's current emotional state. The server uses specialized emotion analysis software to analyze the user's stress level and changes in emotion. The specific operations include extracting voice features, recognizing facial expressions, and estimating the emotional state.
[0187] Step 5:
[0188] The server identifies areas of weakness by comparing extracted text information and emotional states with the user's learning history. The input is text information and emotional states, and the output is the identified areas of weakness. The server identifies areas where the user particularly struggles by comparing them with a database of past learning history. Specifically, this involves database matching, comparison with past performance, and reflection of emotional information.
[0189] Step 6:
[0190] The server generates similar problems and adjusts their difficulty level. The user's areas of difficulty and emotional state are inputs, and the generated problem set is the output. Emotional information is used to provide problems best suited to the user's current situation. Specific operations include problem selection, difficulty adjustment, and problem set generation.
[0191] Step 7:
[0192] The system provides explanations optimized for the user. The input is the explanation for the generated problem, and the output is the content presented to the user. The server selects the format of the explanation based on the user's emotional state and provides it in video or text format. Specifically, it handles the preparation of the explanation content, format selection, and distribution to the user.
[0193] Step 8:
[0194] Users solve similar problems. The results of their solutions become input and are sent to the server as learning progress data. The user's answers are reflected in the next learning plan as an element of progress management. Specifically, this includes answer input, result evaluation, and transmission of progress data.
[0195] Step 9:
[0196] The server records the user's learning results and creates a new learning plan. The user's answers are the input, and the newly adjusted learning plan is the output. A new learning strategy is determined based on past data and current progress. The specific actions involve updating the database, executing the plan generation algorithm, and notifying the user of the new learning strategy.
[0197] (Application Example 2)
[0198] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0199] In conventional learning systems, when a user makes a mistake, the learning experience is uniform and lacks individualization that takes into account the user's emotional state, leading to a decrease in motivation. Furthermore, providing learning materials that disregard the user's psychological state results in insufficient learning efficiency.
[0200] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0201] In this invention, the server includes an image acquisition means for taking and uploading photos of learning problems that the user has answered incorrectly, an image analysis means for analyzing the uploaded images and extracting text data, and an emotion recognition means for monitoring the user's emotional state in real time. This makes it possible to provide a personalized learning experience that corresponds to the user's emotional state and to present similar problems and explanations of appropriate difficulty levels.
[0202] "Image acquisition means" refers to a device or part of a device used to photograph learning questions that a user has answered incorrectly and to upload the image data.
[0203] "Image analysis means" refers to technologies and devices that perform processing to extract text data from uploaded image data, and which utilize optical character recognition technology.
[0204] "Analysis methods" refer to methods and technologies for identifying areas of difficulty for users by referencing their learning history based on extracted text data.
[0205] A "problem generation method" refers to a function or technology that generates similar problems based on the user's areas of weakness and adjusts the difficulty level of the problems while considering the user's emotional state.
[0206] "Explanation provision means" refers to a method or technology for preparing and providing explanations for similar problems that have been generated, in a format that is appropriate to the user's emotional state.
[0207] "Emotion recognition means" refers to devices or configurations for monitoring a user's emotional state in real time and analyzing the data.
[0208] A "feedback provision method" refers to a method or system for recording learning results and creating a new learning plan based on them.
[0209] This invention is a system for efficiently supporting user learning and aims to provide an individualized learning experience that takes into account the user's emotional state.
[0210] The system primarily consists of terminals, servers, and their respective software components. The terminals are equipped with cameras and microphones, and images are acquired when the user takes a picture of a learning problem they answered incorrectly. This image data is uploaded to a cloud server in real time, where image analysis is performed. Text data is extracted from the images using OCR technology. For optical character recognition, services such as Google Cloud Vision can be used.
[0211] Furthermore, the device has a built-in emotion recognition engine that analyzes the user's voice and facial expressions to detect their emotional state in real time. The Affectiva SDK and other tools can be used for emotion recognition. The server identifies the user's areas of difficulty based on the emotional data and uses a problem generation engine to generate similar problems tailored to the user. In this process, machine learning models and generative AI models are utilized, taking into account the user's past learning history and current emotional state.
[0212] The generated problems are accompanied by explanations in an appropriate format, taking into account the user's emotional state. For example, if the user is relaxed, a concise text explanation is provided; if they are judged to be highly focused, detailed video content is presented. This makes it possible to maximize the learning effect.
[0213] As a concrete example, consider a scenario where a user makes a mistake on an English grammar question. The user takes a picture of the question with their device, and simultaneously, an emotion recognition engine detects the user's frustration. Based on this information, the server provides relaxing video content along with simple review questions. In this way, delivering optimal learning content tailored to the user's emotions improves both motivation and efficiency in learning.
[0214] An example of a prompt might be: "Generate appropriate learning content when the user's facial expression data indicates 'frustration.' Lower the difficulty level and include relaxing video content."
[0215] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0216] Step 1:
[0217] The user takes a picture of the problem and the image is acquired. The input is the image of the problem captured by the camera, and the output is the captured image data. The terminal activates the camera at the user's command and takes a picture of the problem sheet. This image is immediately imported into the system.
[0218] Step 2:
[0219] The device uploads the captured image data to a cloud server. The input is the captured image data, and the output is the image data sent to the server. The uploaded images are ready for image analysis.
[0220] Step 3:
[0221] The server extracts text data from uploaded image data. The input is image data, and the output is extracted text data. The server uses optical character recognition (OCR) technology to recognize characters within the image and convert them into text. Through this process, the problem statement is obtained as text data.
[0222] Step 4:
[0223] The device collects and analyzes voice and facial expression data to recognize the user's emotions. The input is the user's voice and facial expression data, and the output is the analyzed emotion data. The emotion recognition engine processes this data in real time and quantifies the user's emotional state.
[0224] Step 5:
[0225] The server identifies the user's weak areas and generates problems based on text and sentiment data. Input is text and sentiment data, and output is similar problems and explanations tailored to the user. The optimal problem set is constructed using an AI model that analyzes past learning history and sentiment data.
[0226] Step 6:
[0227] The server optimizes and provides explanations for the generated set of problems based on the user's emotional state. The input is the user's emotional state and the generated problems, while the output is the adjusted explanation content. The difficulty and format of the explanations are adjusted to improve the user's learning efficiency.
[0228] Step 7:
[0229] The user solves a provided similar problem and sends the results from their device to the server. The input is the user's answer, and the output is training data for feedback. The server analyzes these results and uses them to build the next training plan.
[0230] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0231] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0232] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0233] [Second Embodiment]
[0234] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0235] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0236] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0237] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0238] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0239] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0240] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0241] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0242] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0243] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0244] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0245] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0246] This system consists of a user terminal and a server that performs analysis and data processing. First, the user takes photos of problems they answered incorrectly during their learning process using the terminal's camera function. This terminal has the capability to upload the image data captured by the user directly to the server.
[0247] The server uses OCR (Optical Character Recognition) technology to extract text data from received images for analysis. This data is stored in the server's data storage and used for subsequent analysis. Through this analysis, the server considers information about problems the user has previously answered incorrectly and how they were categorized, identifying areas where the user has particular weaknesses. This provides a foundation for generating appropriate similar problems tailored to the user's learning progress.
[0248] Next, the server generates similar problems tailored to the user's areas of weakness using a problem generation system. These problems are adjusted in difficulty and selected from an extensive problem database. Furthermore, the explanation provision system creates and provides explanations corresponding to the generated problems in text or video format. These explanations are designed to explain complex concepts in an easy-to-understand manner.
[0249] Users can receive these problems and explanations via their devices and review them. Furthermore, the user's problem-solving process and results are transmitted to the server in real time. The server records this information through a feedback system and updates the user's learning history. Based on these results, a new learning plan is generated to assist with future learning.
[0250] By implementing this entire system, users can efficiently learn to overcome their weak areas, resulting in more effective learning outcomes.
[0251] The following describes the processing flow.
[0252] Step 1:
[0253] Users take photos of problems they answered incorrectly during their studies using the camera on their smartphone or tablet. It is important that the problems are clearly and distinctly visible in these photos.
[0254] Step 2:
[0255] The device saves the captured image data to its own storage and uploads it to the server with user identification information attached. The device also checks the communication status to ensure that the data is transmitted correctly.
[0256] Step 3:
[0257] The server receives the uploaded image data and uses OCR technology to extract text data from the image. The server then uses the extracted characters and mathematical formulas to identify the problem statement and structures the information as text data.
[0258] Step 4:
[0259] The server analyzes the text data and compares it with the user's past learning history to identify areas or topics they struggle with. In this process, it utilizes multiple databases to analyze and classify the user's answer patterns.
[0260] Step 5:
[0261] The server generates similar problems using a problem generation mechanism based on identified areas of weakness. This includes a process of selecting the difficulty level and format of the problems, and the appropriate problems are selected dynamically.
[0262] Step 6:
[0263] The server creates explanations for problems and prepares the content in text or video format. The explanations are provided in a way that includes relevant information and examples to aid understanding.
[0264] Step 7:
[0265] The terminal receives similar problems and explanations sent from the server and provides an interface for visualization to the user. The user can then view this information for learning and review.
[0266] Step 8:
[0267] The user inputs the results of solving the provided problems and their learning progress into their device and sends this information back to the server. This input includes data on whether the answers were correct or incorrect.
[0268] Step 9:
[0269] The server evaluates the user's learning progress based on their input data and creates the next learning plan. Based on these results, it prepares new similar problems and review plans to be used in the next learning session.
[0270] (Example 1)
[0271] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0272] Conventional learning support systems have struggled to efficiently identify individual users' areas of weakness and generate and provide appropriate learning problems based on those areas. Furthermore, the lack of immediate feedback on the provided learning problems and the subsequent updates to learning plans led to decreased user learning efficiency.
[0273] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0274] In this invention, the server includes information acquisition means, information analysis means, and evaluation means. This makes it possible to photograph learning tasks that the user has answered incorrectly, analyze the information to extract textual information, and further identify the user's areas of weakness based on the extracted textual information.
[0275] "Information acquisition means" refers to a device or function that allows a user to take a picture of a learning task and transmit that data to a server.
[0276] The "information analysis means" is a device or function that analyzes the received data and provides a technology for accurately extracting character information from image information.
[0277] The "evaluation means" is a device or function that analyzes and identifies the user's weak areas based on the extracted character information.
[0278] The "problem generation means" is a device or function that creates and provides appropriate similar problems according to the user's weak areas.
[0279] The "explanation providing means" is a device or function that prepares a detailed explanation for the generated problem and presents it to the user.
[0280] The "evaluation providing means" is a device or function that records the user's learning progress and creates a new learning plan based on it.
[0281] This invention provides an educational support system for efficiently overcoming the learning problems that the user has made mistakes in. This system is composed of a terminal used by the user and a server that analyzes and processes data.
[0282] The user uses the camera function of the terminal to take a picture of the image of the problem that was wrong during learning. This terminal is equipped with information acquisition means for uploading the image data to the server as it is. For example, a smartphone or a tablet is used as the terminal.
[0283] The server extracts character information from the photographed image using OCR (Optical Character Recognition) technology as the information analysis means. For this process, technologies such as "Tesseract", which is widely used as OCR software, are used.
[0284] Furthermore, the server analyzes the extracted character information using the evaluation means and identifies the user's weak areas. By analyzing past learning data and the correct answer rate, the user's difficult fields are clarified.
[0285] Using the problem generation means, the server creates similar problems based on the user's weak areas using a problem generation AI model. Technologies such as the "OpenAI API" are utilized in the problem generation AI model. This model has the function of adjusting the problems to an appropriate level of difficulty considering the user's learning history.
[0286] For the generated problems, detailed explanations are provided by the explanation providing means. This explanation is provided to the user in text format or video format. Tools such as "Explain Everything" are used for the video explanation, and the user can intuitively deepen their understanding.
[0287] Finally, by utilizing the feedback providing means, the server records the user's learning progress and automatically formulates a new learning plan. This plan optimizes the subsequent learning content based on the user's proficiency.
[0288] As a specific example, consider the case where the user is stuck on a fraction problem in mathematics. In this case, the user takes a photo of the problem and uploads it to the server. The server analyzes the image, identifies the lack of understanding related to fractions, then generates similar practice problems and prepares detailed explanations. As an example of the prompt text input to the problem generation AI model, content such as "Please generate similar problems and explanations for the user to solve fraction problems they are not good at. While presenting specific examples, consider ways to explain complex concepts simply" is used.
[0289] The flow of the specific process in Example 1 will be described using FIG. 11.
[0290] Step 1:
[0291] During the learning process, users take photos of problems they answered incorrectly using their device's camera. A dedicated application is installed on the device, allowing it to directly send image data to a server. The input is the image data of the problems the user has photographed. The output is the uploading of that data to the server. By selecting the appropriate problems and photographing them with the camera, users provide accurate information for the system to analyze.
[0292] Step 2:
[0293] The server extracts text information from received image data using OCR technology. The input is the image data uploaded to the server. The output is the extracted text data. The server uses OCR software such as "Tesseract" to read the text information in the image and convert it into parseable text. This process clarifies the specific content of the questions the user answered incorrectly.
[0294] Step 3:
[0295] The server analyzes the extracted text data to identify the user's areas of weakness. The input is the extracted text data. The output is a list of the user's areas of weakness. Past learning data and accuracy rates are used in the analysis, and statistical methods are employed to identify the user's weak areas. This enables personalized educational support.
[0296] Step 4:
[0297] The server generates similar problems based on the identified areas of weakness. A generative AI model is used for this process. The input is a list of the user's areas of weakness. The output is the generated similar problems. AI technologies such as the "OpenAI API" are utilized to create appropriate tasks that deepen the user's understanding. The generated problems are tailored to the user's specific learning needs.
[0298] Step 5:
[0299] The server prepares explanations for the generated problems. The input is a similar problem that has been generated. The output is a detailed explanation of the problem, which includes text and video explanations. Tools such as "Explain Everything" are used to create easy-to-understand explanations, helping users to understand the problems more deeply.
[0300] Step 6:
[0301] The user reviews the material based on the questions and explanations received on their device. During this process, the user's answers are sent to the server in real time. Input consists of the user's answers and their results. Output is the server-side update of the learning progress data. The device records the user's answer history and provides it to the server as feedback.
[0302] Step 7:
[0303] The server updates the user's learning plan based on the feedback provided. The input is the updated learning progress data. The output is the new learning plan. The plan generated by the server is used in the user's next learning session. This process optimizes continuous learning.
[0304] (Application Example 1)
[0305] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0306] In today's information society, personalized learning support tailored to each user's learning characteristics is required to respond quickly and accurately to individual learning needs. However, existing learning systems have the challenge of not being able to accurately identify a user's weak areas and provide effective learning content. To solve this problem, an advanced system is needed that analyzes the user's learning history in real time and provides personalized feedback and learning plans.
[0307] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means respectively.
[0308] In this invention, the server includes: a data acquisition means for photographing and uploading learning problems that the user has made mistakes in; a data analysis means for analyzing the uploaded image and extracting character data; an information analysis means for identifying the user's weak field from the extracted character data; and a data processing means for generating visually easy-to-learn video content based on the analysis result of the user's weakness and distributing it to the data device. Thereby, learning content suitable for individual users can be provided, and efficient understanding promotion becomes possible.
[0309] The "data acquisition means" is a device equipped with a function for photographing learning problems that the user has made mistakes in and uploading the data to the server.
[0310] The "data analysis means" is a device having a technique for analyzing the uploaded image and extracting character data from the image.
[0311] The "information analysis means" is a device for performing data processing to identify the user's weak field based on the extracted character data.
[0312] The "information generation means" is a device having a function of creating similar problems based on the user's weak field and providing them.
[0313] The "information providing means" is a device equipped with a function for preparing an explanation for the generated similar problems and presenting it to the user
[0314] The "result providing means" is a device for recording the user's learning progress and creating a new learning program based on the data.
[0315] A "data processing means" is a device that generates visual video content based on the analysis results of the user's weaknesses and distributes it to a data device.
[0316] Embodiments of the present invention are systems for improving user learning efficiency. This system consists of a terminal that the user normally uses and a server that is responsible for analysis and data processing.
[0317] The user first takes a picture of the learning problem they answered incorrectly using the device's camera function. This device has a built-in data acquisition mechanism that allows the user to upload the captured image directly to the server. The server then uses OCR (Optical Character Recognition) technology to analyze the uploaded image and extract the text data. Specifically, an OCR tool such as Tesseract is used.
[0318] The server identifies the user's areas of difficulty based on this text data and generates personalized information using information analysis tools. In this process, it selects appropriate similar problems for the user and adjusts their difficulty level. It also prepares explanations based on user-specific information and presents them to the user through information delivery tools.
[0319] Furthermore, the server records learning progress and presents the user with a new learning program using a results delivery system. The user's weaknesses are analyzed, and based on the results, visually easy-to-understand video content is generated. The generated content is distributed to a data device via a data processing system.
[0320] For example, a user who struggles with an English grammar problem will be shown an animated explanation of that grammatical point to help them solve the problem. Through this process, users can effectively engage in learning.
[0321] As a concrete example, here are some examples of prompt statements for a generative AI model.
[0322] "When solving the following grammar problems, identify common mistakes and generate related explanations."
[0323] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0324] Step 1:
[0325] The user takes a picture of the learning problem with their camera and saves it to their device.
[0326] The input is an image of the problem taken by the user. The user uploads this image to the server using a data acquisition method.
[0327] Step 2:
[0328] The device uploads the captured image to the server.
[0329] The input is captured image data, and the image data is sent to the server as output. This gives the server the foundation to receive the image data and begin processing it.
[0330] Step 3:
[0331] The server uses OCR technology to extract text data from the image.
[0332] The input is image data uploaded to the server, and the output is extracted text data. In this process, an OCR tool such as Tesseract is used to convert the text information in the image into digital text.
[0333] Step 4:
[0334] The server analyzes the extracted text data to identify the user's areas of weakness.
[0335] The input is extracted character data, and the output is information about the user's areas of difficulty. The server uses information analysis tools to analyze the patterns shown in the data and determine whether the user is repeatedly making mistakes in a particular area.
[0336] Step 5:
[0337] The server generates similar problems based on the user's areas of weakness.
[0338] The input is information on areas of weakness, and the output is a generated similar problem. The server uses information generation tools to select appropriate problems from a broad problem database and generates similar problems while adjusting the difficulty level.
[0339] Step 6:
[0340] The server prepares explanations for similar problems that have been generated and distributes them to users.
[0341] The input is a generated similar problem, and the output is related explanatory content. The server uses information delivery methods to create explanations in text or video format, and efforts are made to ensure that the explanations are easy for the user to understand.
[0342] Step 7:
[0343] The server records the user's learning progress and creates a new learning program.
[0344] The input is the user's problem-solving history and its analysis results, and the output is a learning program based on the next learning stage. The server maintains the learning record using a results-providing mechanism and presents the user with appropriate feedback and the next learning plan based on it.
[0345] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0346] This invention embodies a system that provides a personalized learning experience that takes into account the user's emotional state when they are learning incorrectly. The system comprises a terminal equipped with an emotion engine, a server that performs analysis processing, and a user interface.
[0347] First, the user takes a picture of the questions they answered incorrectly during the learning process using the device's camera. The device uploads these images to a server. Additionally, a built-in emotion engine monitors the user's voice and facial expressions in real time. From this data, the system detects the user's current emotional state.
[0348] The server analyzes the received image data using OCR technology and extracts the text data for the problem. Based on this data, it compares it with the user's past learning history to identify areas of difficulty. By also considering emotional data in identifying areas of difficulty, it becomes possible to estimate the user's stress level and motivation and present appropriate learning materials.
[0349] In generating similar problems, the server adjusts the difficulty level based on the user's emotional data. For example, if the server determines that the user is stressed, it can start with easier problems and gradually increase the difficulty. This reduces the burden of learning and allows for more effective review.
[0350] The explanations can be flexibly adapted to user sentiment data to support their understanding. For example, short text can be provided when the user is relaxed, while detailed video content can be offered when the user is highly focused.
[0351] Users progress through the provided similar problems and explanations, and upon completing a problem, the results are sent back to the server from their device. The server records these results as feedback and evaluates the user's progress. The newly generated learning plan is used for the next learning session, continuously providing a learning experience optimized for the user's individual needs.
[0352] For example, if a user makes a mistake on a math problem, the emotion engine detects the user's level of fatigue as soon as they upload the image they took. Based on this emotion information, the server incorporates relaxing problems and videos into the review plan. As a result, the user can learn efficiently while reducing stress.
[0353] The following describes the processing flow.
[0354] Step 1:
[0355] The user takes a picture of the question they answered incorrectly using their device's camera. The image should clearly show the question, the answer choices, and the answer the user selected.
[0356] Step 2:
[0357] The device saves captured images to its storage and uploads them to a server along with the user's identification information. Furthermore, it activates an emotion engine to analyze the user's facial expressions and voice in real time and collect emotion data.
[0358] Step 3:
[0359] The server analyzes the received image data using OCR technology and extracts text data from the image. At this time, the question, answer choices, and answer information are extracted as text.
[0360] Step 4:
[0361] The server analyzes the extracted text data, compares it with the user's learning history, and identifies areas of weakness. This analysis reflects the user's past performance and learning patterns.
[0362] Step 5:
[0363] The server generates similar problems while taking emotional data into consideration. If the user's emotional state is relaxed, it selects more difficult problems; if they are stressed, it prioritizes selecting easier problems.
[0364] Step 6:
[0365] The server prepares explanations related to similar problems that have been generated. If the user's level of understanding is deemed high, a detailed explanation is provided; if their concentration is low, a concise explanation is provided.
[0366] Step 7:
[0367] The terminal receives similar problems and explanations sent from the server and presents them to the user. The user can learn by solving the problems and referring to the explanations.
[0368] Step 8:
[0369] Users input the results of solving new problems into their devices and send that data to the server. This input may include accuracy rates and solving speed.
[0370] Step 9:
[0371] The server combines user response data and sentiment data to evaluate learning progress and generate a next learning plan. This plan is tailored to the user's specific learning style and circumstances.
[0372] (Example 2)
[0373] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0374] Learners sometimes abandon problems they got wrong during the learning process. Furthermore, stress and decreased motivation due to monotonous learning hinder learning efficiency. While there is a need to provide learning experiences tailored to each learner's level of understanding and emotional state, current systems struggle to meet these requirements. Therefore, this invention aims to improve learning efficiency and sustainability by providing a flexible and effective learning process that takes the user's emotional state into consideration.
[0375] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0376] In this invention, the server includes data analysis means, emotion analysis means, and task generation means. This makes it possible to provide learners with personalized learning problems and explanations that take into account the user's areas of weakness and emotional state when they make mistakes on a problem.
[0377] "Data acquisition means" refers to a function that allows users to take photos of learning problems they answered incorrectly, collect the image data, and upload it to the server.
[0378] "Data analysis means" refers to a function that analyzes uploaded image data and extracts text information using optical character recognition.
[0379] "Emotional analysis means" refers to a function that analyzes data such as the user's voice tone and facial expressions to evaluate the user's emotional state in real time.
[0380] The "analysis method" is a function that identifies areas of difficulty for the user based on extracted text information and emotional state.
[0381] The "task generation method" is a function that generates similar problems based on identified areas of weakness and emotional states, and provides them to the user.
[0382] "Information provision means" refers to a function that prepares explanations of similar problems that have been generated and presents them in a format that is adjusted according to the user's emotional state.
[0383] A "progress management tool" is a function that records the user's learning progress and creates a new learning plan that will be useful for future learning.
[0384] The present invention is a system for providing a personalized learning experience that takes into account the user's emotional state when the user learns through incorrect problems. This system includes a terminal equipped with emotion analysis capabilities, a server that performs analysis processing, and a user interface.
[0385] First, the user uses the device's camera to take pictures of the questions they answered incorrectly during their studies. The device uploads this captured image data to the server. The device also has a built-in emotion engine that acquires voice and facial expression data to determine the user's emotional state in real time. This emotion data is also sent to the server.
[0386] Next, the server analyzes the received image data using OCR (Optical Character Recognition) technology to extract the text information in question. The OCR technology used here is software that enables accurate analysis of text data. The extracted text data is then compared with the user's past learning history to identify areas where the user struggles. Furthermore, emotional data is taken into consideration during this identification process, allowing the user's stress level and motivation to be inferred. This makes it possible to present the user with learning materials that are optimal for them.
[0387] Furthermore, the server adjusts the difficulty level of the problems based on the user's emotional data. Specifically, if the user is feeling stressed, it starts with easier problems to reduce the burden. Explanations for the generated problems are provided flexibly according to the user's emotional state to aid their understanding. For example, if the user is relaxed, a concise text-based explanation is provided, while if they are highly focused, detailed video content is offered.
[0388] As a concrete example, consider a scenario where a user makes a mistake on a math problem. When the user sends an image taken with their device to the server, the emotion engine detects the user's level of fatigue. Based on this emotional information, the server proposes a review plan that includes relaxing content. For example, it might include simple math problems or visually relaxing content.
[0389] An example of a prompt for a generative AI model is, "Please write a simple explanation of the following math problem, assuming the user is in a relaxed state." This prompt allows for the customization of the explanation provided to the user, thereby supporting improved learning efficiency.
[0390] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0391] Step 1:
[0392] The user takes a picture of the question they answered incorrectly using the camera. The captured image data becomes the input. The device temporarily stores this image data and prepares to upload it to the server. The specific actions here are activating the camera, capturing the image, and saving the data.
[0393] Step 2:
[0394] The device uploads the captured image data to the server. At this time, the image data is input and sent to the server. Once the server receives the image data, data analysis becomes possible. Specifically, this process involves data transfer over the network and confirmation of receipt.
[0395] Step 3:
[0396] The server analyzes the received image data using OCR technology. The input is image data, and the output is extracted text information. The server uses OCR software to extract character information from the image data and convert it into text data. Specifically, processes such as image preprocessing, character recognition, and text conversion are performed.
[0397] Step 4:
[0398] The server processes the user's voice and facial expression data and performs emotion analysis. Voice and facial expression data are inputs, and the output is the user's current emotional state. The server uses specialized emotion analysis software to analyze the user's stress level and changes in emotion. The specific operations include extracting voice features, recognizing facial expressions, and estimating the emotional state.
[0399] Step 5:
[0400] The server identifies areas of weakness by comparing extracted text information and emotional states with the user's learning history. The input is text information and emotional states, and the output is the identified areas of weakness. The server identifies areas where the user particularly struggles by comparing them with a database of past learning history. Specifically, this involves database matching, comparison with past performance, and reflection of emotional information.
[0401] Step 6:
[0402] The server generates similar problems and adjusts their difficulty level. The user's areas of difficulty and emotional state are inputs, and the generated problem set is the output. Emotional information is used to provide problems best suited to the user's current situation. Specific operations include problem selection, difficulty adjustment, and problem set generation.
[0403] Step 7:
[0404] The system provides explanations optimized for the user. The input is the explanation for the generated problem, and the output is the content presented to the user. The server selects the format of the explanation based on the user's emotional state and provides it in video or text format. Specifically, it handles the preparation of the explanation content, format selection, and distribution to the user.
[0405] Step 8:
[0406] Users solve similar problems. The results of their solutions become input and are sent to the server as learning progress data. The user's answers are reflected in the next learning plan as an element of progress management. Specifically, this includes answer input, result evaluation, and transmission of progress data.
[0407] Step 9:
[0408] The server records the user's learning results and creates a new learning plan. The user's answers are the input, and the newly adjusted learning plan is the output. A new learning strategy is determined based on past data and current progress. The specific actions involve updating the database, executing the plan generation algorithm, and notifying the user of the new learning strategy.
[0409] (Application Example 2)
[0410] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0411] In conventional learning systems, when a user makes a mistake, the learning experience is uniform and lacks individualization that takes into account the user's emotional state, leading to a decrease in motivation. Furthermore, providing learning materials that disregard the user's psychological state results in insufficient learning efficiency.
[0412] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0413] In this invention, the server includes an image acquisition means for taking and uploading photos of learning problems that the user has answered incorrectly, an image analysis means for analyzing the uploaded images and extracting text data, and an emotion recognition means for monitoring the user's emotional state in real time. This makes it possible to provide a personalized learning experience that corresponds to the user's emotional state and to present similar problems and explanations of appropriate difficulty levels.
[0414] "Image acquisition means" refers to a device or part of a device used to photograph learning questions that a user has answered incorrectly and to upload the image data.
[0415] "Image analysis means" refers to technologies and devices that perform processing to extract text data from uploaded image data, and which utilize optical character recognition technology.
[0416] "Analysis methods" refer to methods and technologies for identifying areas of difficulty for users by referencing their learning history based on extracted text data.
[0417] A "problem generation method" refers to a function or technology that generates similar problems based on the user's areas of weakness and adjusts the difficulty level of the problems while considering the user's emotional state.
[0418] "Explanation provision means" refers to a method or technology for preparing and providing explanations for similar problems that have been generated, in a format that is appropriate to the user's emotional state.
[0419] "Emotion recognition means" refers to devices or configurations for monitoring a user's emotional state in real time and analyzing the data.
[0420] A "feedback provision method" refers to a method or system for recording learning results and creating a new learning plan based on them.
[0421] This invention is a system for efficiently supporting user learning and aims to provide an individualized learning experience that takes into account the user's emotional state.
[0422] The system primarily consists of terminals, servers, and their respective software components. The terminals are equipped with cameras and microphones, and images are acquired when the user takes a picture of a learning problem they answered incorrectly. This image data is uploaded to a cloud server in real time, where image analysis is performed. Text data is extracted from the images using OCR technology. For optical character recognition, services such as Google Cloud Vision can be used.
[0423] Furthermore, the device has a built-in emotion recognition engine that analyzes the user's voice and facial expressions to detect their emotional state in real time. The Affectiva SDK and other tools can be used for emotion recognition. The server identifies the user's areas of difficulty based on the emotional data and uses a problem generation engine to generate similar problems tailored to the user. In this process, machine learning models and generative AI models are utilized, taking into account the user's past learning history and current emotional state.
[0424] The generated problems are accompanied by explanations in an appropriate format, taking into account the user's emotional state. For example, if the user is relaxed, a concise text explanation is provided; if they are judged to be highly focused, detailed video content is presented. This makes it possible to maximize the learning effect.
[0425] As a concrete example, consider a scenario where a user makes a mistake on an English grammar question. The user takes a picture of the question with their device, and simultaneously, an emotion recognition engine detects the user's frustration. Based on this information, the server provides relaxing video content along with simple review questions. In this way, delivering optimal learning content tailored to the user's emotions improves both motivation and efficiency in learning.
[0426] An example of a prompt might be: "Generate appropriate learning content when the user's facial expression data indicates 'frustration.' Lower the difficulty level and include relaxing video content."
[0427] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0428] Step 1:
[0429] The user takes a picture of the problem and the image is acquired. The input is the image of the problem captured by the camera, and the output is the captured image data. The terminal activates the camera at the user's command and takes a picture of the problem sheet. This image is immediately imported into the system.
[0430] Step 2:
[0431] The device uploads the captured image data to a cloud server. The input is the captured image data, and the output is the image data sent to the server. The uploaded images are ready for image analysis.
[0432] Step 3:
[0433] The server extracts text data from uploaded image data. The input is image data, and the output is extracted text data. The server uses optical character recognition (OCR) technology to recognize characters within the image and convert them into text. Through this process, the problem statement is obtained as text data.
[0434] Step 4:
[0435] The device collects and analyzes voice and facial expression data to recognize the user's emotions. The input is the user's voice and facial expression data, and the output is the analyzed emotion data. The emotion recognition engine processes this data in real time and quantifies the user's emotional state.
[0436] Step 5:
[0437] The server identifies the user's weak areas and generates problems based on text and sentiment data. Input is text and sentiment data, and output is similar problems and explanations tailored to the user. The optimal problem set is constructed using an AI model that analyzes past learning history and sentiment data.
[0438] Step 6:
[0439] The server optimizes and provides explanations for the generated set of problems based on the user's emotional state. The input is the user's emotional state and the generated problems, while the output is the adjusted explanation content. The difficulty and format of the explanations are adjusted to improve the user's learning efficiency.
[0440] Step 7:
[0441] The user solves a provided similar problem and sends the results from their device to the server. The input is the user's answer, and the output is training data for feedback. The server analyzes these results and uses them to build the next training plan.
[0442] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0443] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0444] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0445] [Third Embodiment]
[0446] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0447] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0448] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0449] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0450] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0451] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0452] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0453] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0454] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0455] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0456] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0457] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0458] This system consists of a user terminal and a server that performs analysis and data processing. First, the user takes photos of problems they answered incorrectly during their learning process using the terminal's camera function. This terminal has the capability to upload the image data captured by the user directly to the server.
[0459] The server uses OCR (Optical Character Recognition) technology to extract text data from received images for analysis. This data is stored in the server's data storage and used for subsequent analysis. Through this analysis, the server considers information about problems the user has previously answered incorrectly and how they were categorized, identifying areas where the user has particular weaknesses. This provides a foundation for generating appropriate similar problems tailored to the user's learning progress.
[0460] Next, the server generates similar problems tailored to the user's areas of weakness using a problem generation system. These problems are adjusted in difficulty and selected from an extensive problem database. Furthermore, the explanation provision system creates and provides explanations corresponding to the generated problems in text or video format. These explanations are designed to explain complex concepts in an easy-to-understand manner.
[0461] Users can receive these problems and explanations via their devices and review them. Furthermore, the user's problem-solving process and results are transmitted to the server in real time. The server records this information through a feedback system and updates the user's learning history. Based on these results, a new learning plan is generated to assist with future learning.
[0462] By implementing this entire system, users can efficiently learn to overcome their weak areas, resulting in more effective learning outcomes.
[0463] The following describes the processing flow.
[0464] Step 1:
[0465] Users take photos of problems they answered incorrectly during their studies using the camera on their smartphone or tablet. It is important that the problems are clearly and distinctly visible in these photos.
[0466] Step 2:
[0467] The device saves the captured image data to its own storage and uploads it to the server with user identification information attached. The device also checks the communication status to ensure that the data is transmitted correctly.
[0468] Step 3:
[0469] The server receives the uploaded image data and uses OCR technology to extract text data from the image. The server then uses the extracted characters and mathematical formulas to identify the problem statement and structures the information as text data.
[0470] Step 4:
[0471] The server analyzes the text data and compares it with the user's past learning history to identify areas or topics they struggle with. In this process, it utilizes multiple databases to analyze and classify the user's answer patterns.
[0472] Step 5:
[0473] The server generates similar problems using a problem generation mechanism based on identified areas of weakness. This includes a process of selecting the difficulty level and format of the problems, and the appropriate problems are selected dynamically.
[0474] Step 6:
[0475] The server creates explanations for problems and prepares the content in text or video format. The explanations are provided in a way that includes relevant information and examples to aid understanding.
[0476] Step 7:
[0477] The terminal receives similar problems and explanations sent from the server and provides an interface for visualization to the user. The user can then view this information for learning and review.
[0478] Step 8:
[0479] The user inputs the results of solving the provided problems and their learning progress into their device and sends this information back to the server. This input includes data on whether the answers were correct or incorrect.
[0480] Step 9:
[0481] The server evaluates the user's learning progress based on their input data and creates the next learning plan. Based on these results, it prepares new similar problems and review plans to be used in the next learning session.
[0482] (Example 1)
[0483] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0484] Conventional learning support systems have struggled to efficiently identify individual users' areas of weakness and generate and provide appropriate learning problems based on those areas. Furthermore, the lack of immediate feedback on the provided learning problems and the subsequent updates to learning plans led to decreased user learning efficiency.
[0485] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0486] In this invention, the server includes information acquisition means, information analysis means, and evaluation means. This makes it possible to photograph learning tasks that the user has answered incorrectly, analyze the information to extract textual information, and further identify the user's areas of weakness based on the extracted textual information.
[0487] "Information acquisition means" refers to a device or function that allows a user to take a picture of a learning task and transmit that data to a server.
[0488] "Information analysis means" refers to a device or function that analyzes received data and provides technology to accurately extract textual information from image information.
[0489] "Evaluation means" refers to a device or function that analyzes and identifies areas of weakness for the user based on extracted textual information.
[0490] "Problem generation means" refers to a device or function that creates and provides appropriate similar problems according to the user's areas of weakness.
[0491] "Explanation provision means" refers to a device or function that prepares a detailed explanation for a generated problem and presents it to the user.
[0492] "Evaluation provision means" refers to a device or function that records the user's learning progress and creates a new learning plan based on that progress.
[0493] This invention provides an educational support system that enables users to efficiently overcome learning tasks they have made mistakes on. The system consists of a user terminal and a server that performs data analysis and processing.
[0494] Users use their device's camera function to take pictures of problems they answered incorrectly during their studies. This device is equipped with a means to upload the image data directly to a server. For example, a smartphone or tablet can be used as the device.
[0495] The server extracts text information from captured images using OCR (Optical Character Recognition) technology as an information analysis method. This process utilizes technologies such as "Tesseract," which is widely used as OCR software.
[0496] Furthermore, the server analyzes the extracted text information using evaluation tools to identify the user's weak areas. By analyzing past learning data and accuracy rates, the user's areas of difficulty are clearly defined.
[0497] Using a task generation method, the server creates similar tasks based on the user's areas of weakness using a generation AI model. Technologies such as the "OpenAI API" are utilized in this generation AI model. This model has the functionality to adjust the task difficulty to an appropriate level, taking into account the user's learning history.
[0498] The generated tasks are accompanied by detailed explanations provided through explanatory tools. These explanations are offered to users in either text or video format. Video explanations utilize tools such as "Explain Everything," allowing users to intuitively deepen their understanding.
[0499] Finally, using feedback mechanisms, the server records the user's learning progress and automatically generates a new learning plan. This plan optimizes subsequent learning content based on the user's proficiency level.
[0500] As a concrete example, consider a scenario where a user struggles with a fraction problem in mathematics. In this case, the user takes a picture of the problem and uploads it to the server. The server analyzes the image, identifies the user's lack of understanding related to fractions, and then generates similar practice problems and prepares detailed explanations. An example of a prompt to be input to the generation AI model would be: "Generate similar problems and explanations to help the user solve the fraction problem they struggle with. Please provide concrete examples and consider ways to explain complex concepts simply."
[0501] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0502] Step 1:
[0503] During the learning process, users take photos of problems they answered incorrectly using their device's camera. A dedicated application is installed on the device, allowing it to directly send image data to a server. The input is the image data of the problems the user has photographed. The output is the uploading of that data to the server. By selecting the appropriate problems and photographing them with the camera, users provide accurate information for the system to analyze.
[0504] Step 2:
[0505] The server extracts text information from received image data using OCR technology. The input is the image data uploaded to the server. The output is the extracted text data. The server uses OCR software such as "Tesseract" to read the text information in the image and convert it into parseable text. This process clarifies the specific content of the questions the user answered incorrectly.
[0506] Step 3:
[0507] The server analyzes the extracted text data to identify the user's areas of weakness. The input is the extracted text data. The output is a list of the user's areas of weakness. Past learning data and accuracy rates are used in the analysis, and statistical methods are employed to identify the user's weak areas. This enables personalized educational support.
[0508] Step 4:
[0509] The server generates similar problems based on the identified areas of weakness. A generative AI model is used for this process. The input is a list of the user's areas of weakness. The output is the generated similar problems. AI technologies such as the "OpenAI API" are utilized to create appropriate tasks that deepen the user's understanding. The generated problems are tailored to the user's specific learning needs.
[0510] Step 5:
[0511] The server prepares explanations for the generated problems. The input is a similar problem that has been generated. The output is a detailed explanation of the problem, which includes text and video explanations. Tools such as "Explain Everything" are used to create easy-to-understand explanations, helping users to understand the problems more deeply.
[0512] Step 6:
[0513] The user reviews the material based on the questions and explanations received on their device. During this process, the user's answers are sent to the server in real time. Input consists of the user's answers and their results. Output is the server-side update of the learning progress data. The device records the user's answer history and provides it to the server as feedback.
[0514] Step 7:
[0515] The server updates the user's learning plan based on the feedback provided. The input is the updated learning progress data. The output is the new learning plan. The plan generated by the server is used in the user's next learning session. This process optimizes continuous learning.
[0516] (Application Example 1)
[0517] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0518] In today's information society, personalized learning support tailored to each user's learning characteristics is required to respond quickly and accurately to individual learning needs. However, existing learning systems have the challenge of not being able to accurately identify a user's weak areas and provide effective learning content. To solve this problem, an advanced system is needed that analyzes the user's learning history in real time and provides personalized feedback and learning plans.
[0519] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0520] In this invention, the server includes data acquisition means for taking and uploading photos of learning problems that the user has answered incorrectly; data analysis means for analyzing the uploaded images and extracting text data; information analysis means for identifying the user's weak areas from the extracted text data; and data processing means for generating easy-to-learn visual video content based on the analysis results of the user's weaknesses and distributing it to a data device. This enables the provision of learning content tailored to each individual user and facilitates efficient understanding.
[0521] A "data acquisition means" is a device equipped with the function of taking a picture of a learning problem that the user answered incorrectly and uploading that data to a server.
[0522] A "data analysis means" is a device that has the technology to analyze uploaded images and extract text data from those images.
[0523] An "information analysis tool" is a device that performs data processing based on extracted text data to identify areas where the user struggles.
[0524] An "information generation means" is a device that has the function of creating and providing similar problems based on the user's areas of weakness.
[0525] An "information provision device" is a device equipped with the function of preparing and presenting explanations for similar problems that have been generated to the user.
[0526] A "results-providing device" is a device that records the user's learning progress and creates a new learning program based on that data.
[0527] A "data processing means" is a device that generates visual video content based on the analysis results of the user's weaknesses and distributes it to a data device.
[0528] Embodiments of the present invention are systems for improving user learning efficiency. This system consists of a terminal that the user normally uses and a server that is responsible for analysis and data processing.
[0529] The user first takes a picture of the learning problem they answered incorrectly using the device's camera function. This device has a built-in data acquisition mechanism that allows the user to upload the captured image directly to the server. The server then uses OCR (Optical Character Recognition) technology to analyze the uploaded image and extract the text data. Specifically, an OCR tool such as Tesseract is used.
[0530] The server identifies the user's areas of difficulty based on this text data and generates personalized information using information analysis tools. In this process, it selects appropriate similar problems for the user and adjusts their difficulty level. It also prepares explanations based on user-specific information and presents them to the user through information delivery tools.
[0531] Furthermore, the server records learning progress and presents the user with a new learning program using a results delivery system. The user's weaknesses are analyzed, and based on the results, visually easy-to-understand video content is generated. The generated content is distributed to a data device via a data processing system.
[0532] For example, a user who struggles with an English grammar problem will be shown an animated explanation of that grammatical point to help them solve the problem. Through this process, users can effectively engage in learning.
[0533] As a concrete example, here are some examples of prompt statements for a generative AI model.
[0534] "When solving the following grammar problems, identify common mistakes and generate related explanations."
[0535] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0536] Step 1:
[0537] The user takes a picture of the learning problem with their camera and saves it to their device.
[0538] The input is an image of the problem taken by the user. The user uploads this image to the server using a data acquisition method.
[0539] Step 2:
[0540] The device uploads the captured image to the server.
[0541] The input is captured image data, and the image data is sent to the server as output. This gives the server the foundation to receive the image data and begin processing it.
[0542] Step 3:
[0543] The server uses OCR technology to extract text data from the image.
[0544] The input is image data uploaded to the server, and the output is extracted text data. In this process, an OCR tool such as Tesseract is used to convert the text information in the image into digital text.
[0545] Step 4:
[0546] The server analyzes the extracted text data to identify the user's areas of weakness.
[0547] The input is extracted character data, and the output is information about the user's areas of difficulty. The server uses information analysis tools to analyze the patterns shown in the data and determine whether the user is repeatedly making mistakes in a particular area.
[0548] Step 5:
[0549] The server generates similar problems based on the user's areas of weakness.
[0550] The input is information on areas of weakness, and the output is a generated similar problem. The server uses information generation tools to select appropriate problems from a broad problem database and generates similar problems while adjusting the difficulty level.
[0551] Step 6:
[0552] The server prepares explanations for similar problems that have been generated and distributes them to users.
[0553] The input is a generated similar problem, and the output is related explanatory content. The server uses information delivery methods to create explanations in text or video format, and efforts are made to ensure that the explanations are easy for the user to understand.
[0554] Step 7:
[0555] The server records the user's learning progress and creates a new learning program.
[0556] The input is the user's problem-solving history and its analysis results, and the output is a learning program based on the next learning stage. The server maintains the learning record using a results-providing mechanism and presents the user with appropriate feedback and the next learning plan based on it.
[0557] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0558] This invention embodies a system that provides a personalized learning experience that takes into account the user's emotional state when learning incorrect information. The system comprises a terminal equipped with an emotion engine, a server that performs analytical processing, and a user interface.
[0559] First, the user takes a picture of the questions they answered incorrectly during the learning process using the device's camera. The device uploads these images to a server. Additionally, a built-in emotion engine monitors the user's voice and facial expressions in real time. From this data, the system detects the user's current emotional state.
[0560] The server analyzes the received image data using OCR technology and extracts the text data for the problem. Based on this data, it compares it with the user's past learning history to identify areas of difficulty. By also considering emotional data in identifying areas of difficulty, it becomes possible to estimate the user's stress level and motivation and present appropriate learning materials.
[0561] In generating similar problems, the server adjusts the difficulty level based on the user's emotional data. For example, if the server determines that the user is stressed, it can start with easier problems and gradually increase the difficulty. This reduces the burden of learning and allows for more effective review.
[0562] The explanations can be flexibly adapted to user sentiment data to support their understanding. For example, short text can be provided when the user is relaxed, while detailed video content can be offered when the user is highly focused.
[0563] Users progress through the provided similar problems and explanations, and upon completing a problem, the results are sent back to the server from their device. The server records these results as feedback and evaluates the user's progress. The newly generated learning plan is used for the next learning session, continuously providing a learning experience optimized for the user's individual needs.
[0564] For example, if a user makes a mistake on a math problem, the emotion engine detects the user's level of fatigue as soon as they upload the image they took. Based on this emotion information, the server incorporates relaxing problems and videos into the review plan. As a result, the user can learn efficiently while reducing stress.
[0565] The following describes the processing flow.
[0566] Step 1:
[0567] The user takes a picture of the question they answered incorrectly using their device's camera. The image should clearly show the question, the answer choices, and the answer the user selected.
[0568] Step 2:
[0569] The device saves captured images to its storage and uploads them to a server along with the user's identification information. Furthermore, it activates an emotion engine to analyze the user's facial expressions and voice in real time and collect emotion data.
[0570] Step 3:
[0571] The server analyzes the received image data using OCR technology and extracts text data from the image. At this time, the question, answer choices, and answer information are extracted as text.
[0572] Step 4:
[0573] The server analyzes the extracted text data, compares it with the user's learning history, and identifies areas of weakness. This analysis reflects the user's past performance and learning patterns.
[0574] Step 5:
[0575] The server generates similar problems while taking emotional data into consideration. If the user's emotional state is relaxed, it selects more difficult problems; if they are stressed, it prioritizes selecting easier problems.
[0576] Step 6:
[0577] The server prepares explanations related to similar problems that have been generated. If the user's level of understanding is deemed high, a detailed explanation is provided; if their concentration is low, a concise explanation is provided.
[0578] Step 7:
[0579] The terminal receives similar problems and explanations sent from the server and presents them to the user. The user can learn by solving the problems and referring to the explanations.
[0580] Step 8:
[0581] Users input the results of solving new problems into their devices and send that data to the server. This input may include accuracy rates and solving speed.
[0582] Step 9:
[0583] The server combines user response data and sentiment data to evaluate learning progress and generate a next learning plan. This plan is tailored to the user's specific learning style and circumstances.
[0584] (Example 2)
[0585] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0586] Learners sometimes abandon problems they got wrong during the learning process. Furthermore, stress and decreased motivation due to monotonous learning hinder learning efficiency. While there is a need to provide learning experiences tailored to each learner's level of understanding and emotional state, current systems struggle to meet these requirements. Therefore, this invention aims to improve learning efficiency and sustainability by providing a flexible and effective learning process that takes the user's emotional state into consideration.
[0587] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0588] In this invention, the server includes data analysis means, emotion analysis means, and task generation means. This makes it possible to provide learners with personalized learning problems and explanations that take into account the user's areas of weakness and emotional state when they make mistakes on a problem.
[0589] "Data acquisition means" refers to a function that allows users to take photos of learning problems they answered incorrectly, collect the image data, and upload it to the server.
[0590] "Data analysis means" refers to a function that analyzes uploaded image data and extracts text information using optical character recognition.
[0591] "Emotional analysis means" refers to a function that analyzes data such as the user's voice tone and facial expressions to evaluate the user's emotional state in real time.
[0592] The "analysis method" is a function that identifies areas of difficulty for the user based on extracted text information and emotional state.
[0593] The "task generation method" is a function that generates similar problems based on identified areas of weakness and emotional states, and provides them to the user.
[0594] "Information provision means" refers to a function that prepares explanations of similar problems that have been generated and presents them in a format that is adjusted according to the user's emotional state.
[0595] A "progress management tool" is a function that records the user's learning progress and creates a new learning plan that will be useful for future learning.
[0596] The present invention is a system for providing a personalized learning experience that takes into account the user's emotional state when the user learns through incorrect problems. This system includes a terminal equipped with emotion analysis capabilities, a server that performs analysis processing, and a user interface.
[0597] First, the user uses the device's camera to take pictures of the questions they answered incorrectly during their studies. The device uploads this captured image data to the server. The device also has a built-in emotion engine that acquires voice and facial expression data to determine the user's emotional state in real time. This emotion data is also sent to the server.
[0598] Next, the server analyzes the received image data using OCR (Optical Character Recognition) technology to extract the text information in question. The OCR technology used here is software that enables accurate analysis of text data. The extracted text data is then compared with the user's past learning history to identify areas where the user struggles. Furthermore, emotional data is taken into consideration during this identification process, allowing the user's stress level and motivation to be inferred. This makes it possible to present the user with learning materials that are optimal for them.
[0599] Furthermore, the server adjusts the difficulty level of the problems based on the user's emotional data. Specifically, if the user is feeling stressed, it starts with easier problems to reduce the burden. Explanations for the generated problems are provided flexibly according to the user's emotional state to aid their understanding. For example, if the user is relaxed, a concise text-based explanation is provided, while if they are highly focused, detailed video content is offered.
[0600] As a concrete example, consider a scenario where a user makes a mistake on a math problem. When the user sends an image taken with their device to the server, the emotion engine detects the user's level of fatigue. Based on this emotional information, the server proposes a review plan that includes relaxing content. For example, it might include simple math problems or visually relaxing content.
[0601] An example of a prompt for a generative AI model is, "Write a simple explanation for the following math problem, assuming the user is in a relaxed state." This prompt allows for the customization of the explanation provided to the user, supporting improved learning efficiency.
[0602] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0603] Step 1:
[0604] The user takes a picture of the question they answered incorrectly using the camera. The captured image data becomes the input. The device temporarily stores this image data and prepares to upload it to the server. The specific actions here are activating the camera, capturing the image, and saving the data.
[0605] Step 2:
[0606] The device uploads the captured image data to the server. At this time, the image data is input and sent to the server. Once the server receives the image data, data analysis becomes possible. Specifically, this process involves data transfer over the network and confirmation of receipt.
[0607] Step 3:
[0608] The server analyzes the received image data using OCR technology. The input is image data, and the output is extracted text information. The server uses OCR software to extract character information from the image data and convert it into text data. Specifically, processes such as image preprocessing, character recognition, and text conversion are performed.
[0609] Step 4:
[0610] The server processes the user's voice and facial expression data and performs emotion analysis. Voice and facial expression data are inputs, and the output is the user's current emotional state. The server uses specialized emotion analysis software to analyze the user's stress level and changes in emotion. The specific operations include extracting voice features, recognizing facial expressions, and estimating the emotional state.
[0611] Step 5:
[0612] The server identifies areas of weakness by comparing extracted text information and emotional states with the user's learning history. The input is text information and emotional states, and the output is the identified areas of weakness. The server identifies areas where the user particularly struggles by comparing them with a database of past learning history. Specifically, this involves database matching, comparison with past performance, and reflection of emotional information.
[0613] Step 6:
[0614] The server generates similar problems and adjusts their difficulty level. The user's areas of difficulty and emotional state are inputs, and the generated problem set is the output. Emotional information is used to provide problems best suited to the user's current situation. Specific operations include problem selection, difficulty adjustment, and problem set generation.
[0615] Step 7:
[0616] The system provides explanations optimized for the user. The input is the explanation for the generated problem, and the output is the content presented to the user. The server selects the format of the explanation based on the user's emotional state and provides it in video or text format. Specifically, it handles the preparation of the explanation content, format selection, and distribution to the user.
[0617] Step 8:
[0618] Users solve similar problems. The results of their solutions become input and are sent to the server as learning progress data. The user's answers are reflected in the next learning plan as an element of progress management. Specifically, this includes answer input, result evaluation, and transmission of progress data.
[0619] Step 9:
[0620] The server records the user's learning results and creates a new learning plan. The user's answers are the input, and the newly adjusted learning plan is the output. A new learning strategy is determined based on past data and current progress. The specific actions involve updating the database, executing the plan generation algorithm, and notifying the user of the new learning strategy.
[0621] (Application Example 2)
[0622] Next, we will explain Application Example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0623] In conventional learning systems, when a user makes a mistake, the learning experience is uniform and lacks individualization that takes into account the user's emotional state, leading to a decrease in motivation. Furthermore, providing learning materials that disregard the user's psychological state results in insufficient learning efficiency.
[0624] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0625] In this invention, the server includes an image acquisition means for taking and uploading photos of learning problems that the user has answered incorrectly, an image analysis means for analyzing the uploaded images and extracting text data, and an emotion recognition means for monitoring the user's emotional state in real time. This makes it possible to provide a personalized learning experience that corresponds to the user's emotional state and to present similar problems and explanations of appropriate difficulty levels.
[0626] "Image acquisition means" refers to a device or part of a device used to photograph learning questions that a user has answered incorrectly and to upload the image data.
[0627] "Image analysis means" refers to technologies and devices that perform processing to extract text data from uploaded image data, and which utilize optical character recognition technology.
[0628] "Analysis methods" refer to methods and technologies for identifying areas of difficulty for users by referencing their learning history based on extracted text data.
[0629] A "problem generation method" refers to a function or technology that generates similar problems based on the user's areas of weakness and adjusts the difficulty level of the problems while considering the user's emotional state.
[0630] "Explanation provision means" refers to a method or technology for preparing and providing explanations for similar problems that have been generated, in a format that is appropriate to the user's emotional state.
[0631] "Emotion recognition means" refers to devices or configurations for monitoring a user's emotional state in real time and analyzing the data.
[0632] A "feedback provision method" refers to a method or system for recording learning results and creating a new learning plan based on them.
[0633] This invention is a system for efficiently supporting user learning and aims to provide an individualized learning experience that takes into account the user's emotional state.
[0634] The system primarily consists of terminals, servers, and their respective software components. The terminals are equipped with cameras and microphones, and images are acquired when the user takes a picture of a learning problem they answered incorrectly. This image data is uploaded to a cloud server in real time, where image analysis is performed. Text data is extracted from the images using OCR technology. For optical character recognition, services such as Google Cloud Vision can be used.
[0635] Furthermore, the device has a built-in emotion recognition engine that analyzes the user's voice and facial expressions to detect their emotional state in real time. The Affectiva SDK and other tools can be used for emotion recognition. The server identifies the user's areas of difficulty based on the emotional data and uses a problem generation engine to generate similar problems tailored to the user. In this process, machine learning models and generative AI models are utilized, taking into account the user's past learning history and current emotional state.
[0636] The generated problems are accompanied by explanations in an appropriate format, taking into account the user's emotional state. For example, if the user is relaxed, a concise text explanation is provided; if they are judged to be highly focused, detailed video content is presented. This makes it possible to maximize the learning effect.
[0637] As a concrete example, consider a scenario where a user makes a mistake on an English grammar question. The user takes a picture of the question with their device, and simultaneously, an emotion recognition engine detects the user's frustration. Based on this information, the server provides relaxing video content along with simple review questions. In this way, delivering optimal learning content tailored to the user's emotions improves both motivation and efficiency in learning.
[0638] An example of a prompt might be: "Generate appropriate learning content when the user's facial expression data indicates 'frustration.' Lower the difficulty level and include relaxing video content."
[0639] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0640] Step 1:
[0641] The user takes a picture of the problem and the image is acquired. The input is the image of the problem captured by the camera, and the output is the captured image data. The terminal activates the camera at the user's command and takes a picture of the problem sheet. This image is immediately imported into the system.
[0642] Step 2:
[0643] The device uploads the captured image data to a cloud server. The input is the captured image data, and the output is the image data sent to the server. The uploaded images are ready for image analysis.
[0644] Step 3:
[0645] The server extracts text data from uploaded image data. The input is image data, and the output is extracted text data. The server uses optical character recognition (OCR) technology to recognize characters within the image and convert them into text. Through this process, the problem statement is obtained as text data.
[0646] Step 4:
[0647] The device collects and analyzes voice and facial expression data to recognize the user's emotions. The input is the user's voice and facial expression data, and the output is the analyzed emotion data. The emotion recognition engine processes this data in real time and quantifies the user's emotional state.
[0648] Step 5:
[0649] The server identifies the user's weak areas and generates problems based on text data and sentiment data. Input is text data and sentiment data, and output is similar problems and explanations tailored to the user. The optimal problem set is constructed using an AI model that analyzes past learning history and sentiment data.
[0650] Step 6:
[0651] The server optimizes and provides explanations for the generated set of problems based on the user's emotional state. The input is the user's emotional state and the generated problems, while the output is the adjusted explanatory content. The difficulty and format of the explanations are adjusted to improve the user's learning efficiency.
[0652] Step 7:
[0653] The user solves a provided similar problem and sends the results from their device to the server. The input is the user's answer, and the output is training data for feedback. The server analyzes these results and uses them to build the next training plan.
[0654] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0655] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0656] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0657] [Fourth Embodiment]
[0658] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0659] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0660] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0661] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0662] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0663] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0664] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0665] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0666] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0667] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0668] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0669] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0670] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0671] This system consists of a user terminal and a server that performs analysis and data processing. First, the user takes photos of problems they answered incorrectly during their learning process using the terminal's camera function. This terminal has the capability to upload the image data captured by the user directly to the server.
[0672] The server uses OCR (Optical Character Recognition) technology to extract text data from received images for analysis. This data is stored in the server's data storage and used for subsequent analysis. Through this analysis, the server considers information about problems the user has previously answered incorrectly and how they were categorized, identifying areas where the user has particular weaknesses. This provides a foundation for generating appropriate similar problems tailored to the user's learning progress.
[0673] Next, the server generates similar problems tailored to the user's areas of weakness using a problem generation system. These problems are adjusted in difficulty and selected from an extensive problem database. Furthermore, the explanation provision system creates and provides explanations corresponding to the generated problems in text or video format. These explanations are designed to explain complex concepts in an easy-to-understand manner.
[0674] Users can receive these problems and explanations via their devices and review them. Furthermore, the user's problem-solving process and results are transmitted to the server in real time. The server records this information through a feedback system and updates the user's learning history. Based on these results, a new learning plan is generated to assist with future learning.
[0675] By implementing this entire system, users can efficiently learn to overcome their weak areas, resulting in more effective learning outcomes.
[0676] The following describes the processing flow.
[0677] Step 1:
[0678] Users take photos of problems they answered incorrectly during their studies using the camera on their smartphone or tablet. It is important that the problems are clearly and distinctly visible in these photos.
[0679] Step 2:
[0680] The device saves the captured image data to its own storage and uploads it to the server with user identification information attached. The device also checks the communication status to ensure that the data is transmitted correctly.
[0681] Step 3:
[0682] The server receives the uploaded image data and uses OCR technology to extract text data from the image. The server then uses the extracted characters and mathematical formulas to identify the problem statement and structures the information as text data.
[0683] Step 4:
[0684] The server analyzes the text data and compares it with the user's past learning history to identify areas or topics they struggle with. In this process, it utilizes multiple databases to analyze and classify the user's answer patterns.
[0685] Step 5:
[0686] The server generates similar problems using a problem generation mechanism based on identified areas of weakness. This includes a process of selecting the difficulty level and format of the problems, and the appropriate problems are selected dynamically.
[0687] Step 6:
[0688] The server creates explanations for problems and prepares the content in text or video format. The explanations are provided in a way that includes relevant information and examples to aid understanding.
[0689] Step 7:
[0690] The terminal receives similar problems and explanations sent from the server and provides an interface for visualization to the user. The user can then view this information for learning and review.
[0691] Step 8:
[0692] The user inputs the results of solving the provided problems and their learning progress into their device and sends this information back to the server. This input includes data on whether the answers were correct or incorrect.
[0693] Step 9:
[0694] The server evaluates the user's learning progress based on their input data and creates the next learning plan. Based on these results, it prepares new similar problems and review plans to be used in the next learning session.
[0695] (Example 1)
[0696] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0697] Conventional learning support systems have struggled to efficiently identify individual users' areas of weakness and generate and provide appropriate learning problems based on those areas. Furthermore, the lack of immediate feedback on the provided learning problems and the subsequent updates to learning plans led to decreased user learning efficiency.
[0698] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0699] In this invention, the server includes information acquisition means, information analysis means, and evaluation means. This makes it possible to photograph learning tasks that the user has answered incorrectly, analyze the information to extract textual information, and further identify the user's areas of weakness based on the extracted textual information.
[0700] "Information acquisition means" refers to a device or function that allows a user to take a picture of a learning task and transmit that data to a server.
[0701] "Information analysis means" refers to a device or function that analyzes received data and provides technology to accurately extract textual information from image information.
[0702] "Evaluation means" refers to a device or function that analyzes and identifies areas of weakness for the user based on extracted textual information.
[0703] "Problem generation means" refers to a device or function that creates and provides appropriate similar problems according to the user's areas of weakness.
[0704] "Explanation provision means" refers to a device or function that prepares a detailed explanation for a generated problem and presents it to the user.
[0705] "Evaluation provision means" refers to a device or function that records the user's learning progress and creates a new learning plan based on that progress.
[0706] This invention provides an educational support system that enables users to efficiently overcome learning tasks they have made mistakes on. The system consists of a user terminal and a server that performs data analysis and processing.
[0707] Users use their device's camera function to take pictures of problems they answered incorrectly during their studies. This device is equipped with a means to upload the image data directly to a server. For example, a smartphone or tablet can be used as the device.
[0708] The server extracts text information from captured images using OCR (Optical Character Recognition) technology as an information analysis method. This process utilizes technologies such as "Tesseract," which is widely used as OCR software.
[0709] Furthermore, the server analyzes the extracted text information using evaluation tools to identify the user's weak areas. By analyzing past learning data and accuracy rates, the user's areas of difficulty are clearly defined.
[0710] Using a task generation method, the server creates similar tasks based on the user's areas of weakness using a generation AI model. Technologies such as the "OpenAI API" are utilized in this generation AI model. This model has the functionality to adjust the task difficulty to an appropriate level, taking into account the user's learning history.
[0711] The generated tasks are accompanied by detailed explanations provided through explanatory tools. These explanations are offered to users in either text or video format. Video explanations utilize tools such as "Explain Everything," allowing users to intuitively deepen their understanding.
[0712] Finally, using feedback mechanisms, the server records the user's learning progress and automatically generates a new learning plan. This plan optimizes subsequent learning content based on the user's proficiency level.
[0713] As a concrete example, consider a scenario where a user struggles with a fraction problem in mathematics. In this case, the user takes a picture of the problem and uploads it to the server. The server analyzes the image, identifies the user's lack of understanding related to fractions, and then generates similar practice problems and prepares detailed explanations. An example of a prompt to be input to the generation AI model would be: "Generate similar problems and explanations to help the user solve the fraction problem they struggle with. Please provide concrete examples and consider ways to explain complex concepts simply."
[0714] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0715] Step 1:
[0716] During the learning process, users take photos of problems they answered incorrectly using their device's camera. A dedicated application is installed on the device, allowing it to directly send image data to a server. The input is the image data of the problems the user has photographed. The output is the uploading of that data to the server. By selecting the appropriate problems and photographing them with the camera, users provide accurate information for the system to analyze.
[0717] Step 2:
[0718] The server extracts text information from received image data using OCR technology. The input is the image data uploaded to the server. The output is the extracted text data. The server uses OCR software such as "Tesseract" to read the text information in the image and convert it into parseable text. This process clarifies the specific content of the questions the user answered incorrectly.
[0719] Step 3:
[0720] The server analyzes the extracted text data to identify the user's areas of weakness. The input is the extracted text data. The output is a list of the user's areas of weakness. Past learning data and accuracy rates are used in the analysis, and statistical methods are employed to identify the user's weak areas. This enables personalized educational support.
[0721] Step 4:
[0722] The server generates similar problems based on the identified areas of weakness. A generative AI model is used for this process. The input is a list of the user's areas of weakness. The output is the generated similar problems. AI technologies such as the "OpenAI API" are utilized to create appropriate tasks that deepen the user's understanding. The generated problems are tailored to the user's specific learning needs.
[0723] Step 5:
[0724] The server prepares explanations for the generated problems. The input is a similar problem that has been generated. The output is a detailed explanation of the problem, which includes text and video explanations. Tools such as "Explain Everything" are used to create easy-to-understand explanations, helping users to understand the problems more deeply.
[0725] Step 6:
[0726] The user reviews the material based on the questions and explanations received on their device. During this process, the user's answers are sent to the server in real time. Input consists of the user's answers and their results. Output is the server-side update of the learning progress data. The device records the user's answer history and provides it to the server as feedback.
[0727] Step 7:
[0728] The server updates the user's learning plan based on the feedback provided. The input is the updated learning progress data. The output is the new learning plan. The plan generated by the server is used in the user's next learning session. This process optimizes continuous learning.
[0729] (Application Example 1)
[0730] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0731] In today's information society, personalized learning support tailored to each user's learning characteristics is required to respond quickly and accurately to individual learning needs. However, existing learning systems have the challenge of not being able to accurately identify a user's weak areas and provide effective learning content. To solve this problem, an advanced system is needed that analyzes the user's learning history in real time and provides personalized feedback and learning plans.
[0732] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0733] In this invention, the server includes data acquisition means for taking and uploading photos of learning problems that the user has answered incorrectly; data analysis means for analyzing the uploaded images and extracting text data; information analysis means for identifying the user's weak areas from the extracted text data; and data processing means for generating easy-to-learn visual video content based on the analysis results of the user's weaknesses and distributing it to a data device. This enables the provision of learning content tailored to each individual user and facilitates efficient understanding.
[0734] A "data acquisition means" is a device equipped with the function of taking a picture of a learning problem that the user answered incorrectly and uploading that data to a server.
[0735] A "data analysis means" is a device that has the technology to analyze uploaded images and extract text data from those images.
[0736] An "information analysis tool" is a device that performs data processing based on extracted text data to identify areas where the user struggles.
[0737] An "information generation means" is a device that has the function of creating and providing similar problems based on the user's areas of weakness.
[0738] An "information provision device" is a device equipped with the function of preparing and presenting explanations for similar problems that have been generated to the user.
[0739] A "results-providing device" is a device that records the user's learning progress and creates a new learning program based on that data.
[0740] A "data processing means" is a device that generates visual video content based on the analysis results of the user's weaknesses and distributes it to a data device.
[0741] Embodiments of the present invention are systems for improving user learning efficiency. This system consists of a terminal that the user normally uses and a server that is responsible for analysis and data processing.
[0742] The user first takes a picture of the learning problem they answered incorrectly using the device's camera function. This device has a built-in data acquisition mechanism that allows the user to upload the captured image directly to the server. The server then uses OCR (Optical Character Recognition) technology to analyze the uploaded image and extract the text data. Specifically, an OCR tool such as Tesseract is used.
[0743] The server identifies the user's areas of difficulty based on this text data and generates personalized information using information analysis tools. In this process, it selects appropriate similar problems for the user and adjusts their difficulty level. It also prepares explanations based on user-specific information and presents them to the user through information delivery tools.
[0744] Furthermore, the server records learning progress and presents the user with a new learning program using a results delivery system. The user's weaknesses are analyzed, and based on the results, visually easy-to-understand video content is generated. The generated content is distributed to a data device via a data processing system.
[0745] For example, a user who struggles with an English grammar problem will be shown an animated explanation of that grammatical point to help them solve the problem. Through this process, users can effectively engage in learning.
[0746] As a concrete example, here are some examples of prompt statements for a generative AI model.
[0747] "When solving the following grammar problems, identify common mistakes and generate related explanations."
[0748] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0749] Step 1:
[0750] The user takes a picture of the learning problem with their camera and saves it to their device.
[0751] The input is an image of the problem taken by the user. The user uploads this image to the server using a data acquisition method.
[0752] Step 2:
[0753] The device uploads the captured image to the server.
[0754] The input is captured image data, and the image data is sent to the server as output. This gives the server the foundation to receive the image data and begin processing it.
[0755] Step 3:
[0756] The server uses OCR technology to extract text data from the image.
[0757] The input is image data uploaded to the server, and the output is extracted text data. In this process, an OCR tool such as Tesseract is used to convert the text information in the image into digital text.
[0758] Step 4:
[0759] The server analyzes the extracted text data to identify the user's areas of weakness.
[0760] The input is extracted character data, and the output is information about the user's areas of difficulty. The server uses information analysis tools to analyze the patterns shown in the data and determine whether the user is repeatedly making mistakes in a particular area.
[0761] Step 5:
[0762] The server generates similar problems based on the user's areas of weakness.
[0763] The input is information on areas of weakness, and the output is a generated similar problem. The server uses information generation tools to select appropriate problems from a broad problem database and generates similar problems while adjusting the difficulty level.
[0764] Step 6:
[0765] The server prepares explanations for similar problems that have been generated and distributes them to users.
[0766] The input is a generated similar problem, and the output is related explanatory content. The server uses information delivery methods to create explanations in text or video format, and efforts are made to ensure that the explanations are easy for the user to understand.
[0767] Step 7:
[0768] The server records the user's learning progress and creates a new learning program.
[0769] The input is the user's problem-solving history and its analysis results, and the output is a learning program based on the next learning stage. The server maintains the learning record using a results-providing mechanism and presents the user with appropriate feedback and the next learning plan based on it.
[0770] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0771] This invention embodies a system that provides a personalized learning experience that takes into account the user's emotional state when learning incorrect information. The system comprises a terminal equipped with an emotion engine, a server that performs analytical processing, and a user interface.
[0772] First, the user takes a picture of the questions they answered incorrectly during the learning process using the device's camera. The device uploads these images to a server. Additionally, a built-in emotion engine monitors the user's voice and facial expressions in real time. From this data, the system detects the user's current emotional state.
[0773] The server analyzes the received image data using OCR technology and extracts the text data for the problem. Based on this data, it compares it with the user's past learning history to identify areas of difficulty. By also considering emotional data in identifying areas of difficulty, it becomes possible to estimate the user's stress level and motivation and present appropriate learning materials.
[0774] In generating similar problems, the server adjusts the difficulty level based on the user's emotional data. For example, if the server determines that the user is stressed, it can start with easier problems and gradually increase the difficulty. This reduces the burden of learning and allows for more effective review.
[0775] The explanations can be flexibly adapted to user sentiment data to support their understanding. For example, short text can be provided when the user is relaxed, while detailed video content can be offered when the user is highly focused.
[0776] Users progress through the provided similar problems and explanations, and upon completing a problem, the results are sent back to the server from their device. The server records these results as feedback and evaluates the user's progress. The newly generated learning plan is used for the next learning session, continuously providing a learning experience optimized for the user's individual needs.
[0777] For example, if a user makes a mistake on a math problem, the emotion engine detects the user's level of fatigue as soon as they upload the image they took. Based on this emotion information, the server incorporates relaxing problems and videos into the review plan. As a result, the user can learn efficiently while reducing stress.
[0778] The following describes the processing flow.
[0779] Step 1:
[0780] The user takes a picture of the question they answered incorrectly using their device's camera. The image should clearly show the question, the answer choices, and the answer the user selected.
[0781] Step 2:
[0782] The device saves captured images to its storage and uploads them to a server along with the user's identification information. Furthermore, it activates an emotion engine to analyze the user's facial expressions and voice in real time and collect emotion data.
[0783] Step 3:
[0784] The server analyzes the received image data using OCR technology and extracts text data from the image. At this time, the question, answer choices, and answer information are extracted as text.
[0785] Step 4:
[0786] The server analyzes the extracted text data, compares it with the user's learning history, and identifies areas of weakness. This analysis reflects the user's past performance and learning patterns.
[0787] Step 5:
[0788] The server generates similar problems while taking emotional data into consideration. If the user's emotional state is relaxed, it selects more difficult problems; if they are stressed, it prioritizes selecting easier problems.
[0789] Step 6:
[0790] The server prepares explanations related to similar problems that have been generated. If the user's level of understanding is deemed high, a detailed explanation is provided; if their concentration is low, a concise explanation is provided.
[0791] Step 7:
[0792] The terminal receives similar problems and explanations sent from the server and presents them to the user. The user can learn by solving the problems and referring to the explanations.
[0793] Step 8:
[0794] Users input the results of solving new problems into their devices and send that data to the server. This input may include accuracy rates and solving speed.
[0795] Step 9:
[0796] The server combines user response data and sentiment data to evaluate learning progress and generate a next learning plan. This plan is tailored to the user's specific learning style and circumstances.
[0797] (Example 2)
[0798] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0799] Learners sometimes abandon problems they got wrong during the learning process. Furthermore, stress and decreased motivation due to monotonous learning hinder learning efficiency. While there is a need to provide learning experiences tailored to each learner's level of understanding and emotional state, current systems struggle to meet these requirements. Therefore, this invention aims to improve learning efficiency and sustainability by providing a flexible and effective learning process that takes the user's emotional state into consideration.
[0800] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0801] In this invention, the server includes data analysis means, emotion analysis means, and task generation means. This makes it possible to provide learners with personalized learning problems and explanations that take into account the user's areas of weakness and emotional state when they make mistakes on a problem.
[0802] "Data acquisition means" refers to a function that allows users to take photos of learning problems they answered incorrectly, collect the image data, and upload it to the server.
[0803] "Data analysis means" refers to a function that analyzes uploaded image data and extracts text information using optical character recognition.
[0804] "Emotional analysis means" refers to a function that analyzes data such as the user's voice tone and facial expressions to evaluate the user's emotional state in real time.
[0805] The "analysis method" is a function that identifies areas of difficulty for the user based on extracted text information and emotional state.
[0806] The "task generation method" is a function that generates similar problems based on identified areas of weakness and emotional states, and provides them to the user.
[0807] "Information provision means" refers to a function that prepares explanations of similar problems that have been generated and presents them in a format that is adjusted according to the user's emotional state.
[0808] A "progress management tool" is a function that records the user's learning progress and creates a new learning plan that will be useful for future learning.
[0809] The present invention is a system for providing a personalized learning experience that takes into account the user's emotional state when the user learns through incorrect problems. This system includes a terminal equipped with emotion analysis capabilities, a server that performs analysis processing, and a user interface.
[0810] First, the user uses the device's camera to take pictures of the questions they answered incorrectly during their studies. The device uploads this captured image data to the server. The device also has a built-in emotion engine that acquires voice and facial expression data to determine the user's emotional state in real time. This emotion data is also sent to the server.
[0811] Next, the server analyzes the received image data using OCR (Optical Character Recognition) technology to extract the text information in question. The OCR technology used here is software that enables accurate analysis of text data. The extracted text data is then compared with the user's past learning history to identify areas where the user struggles. Furthermore, emotional data is taken into consideration during this identification process, allowing the user's stress level and motivation to be inferred. This makes it possible to present the user with learning materials that are optimal for them.
[0812] Furthermore, the server adjusts the difficulty level of the problems based on the user's emotional data. Specifically, if the user is feeling stressed, it starts with easier problems to reduce the burden. Explanations for the generated problems are provided flexibly according to the user's emotional state to aid their understanding. For example, if the user is relaxed, a concise text-based explanation is provided, while if they are highly focused, detailed video content is offered.
[0813] As a concrete example, consider a scenario where a user makes a mistake on a math problem. When the user sends an image taken with their device to the server, the emotion engine detects the user's level of fatigue. Based on this emotional information, the server proposes a review plan that includes relaxing content. For example, it might include simple math problems or visually relaxing content.
[0814] An example of a prompt for a generative AI model is, "Write a simple explanation for the following math problem, assuming the user is in a relaxed state." This prompt allows for the customization of the explanation provided to the user, supporting improved learning efficiency.
[0815] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0816] Step 1:
[0817] The user takes a picture of the question they answered incorrectly using the camera. The captured image data becomes the input. The device temporarily stores this image data and prepares to upload it to the server. The specific actions here are activating the camera, capturing the image, and saving the data.
[0818] Step 2:
[0819] The device uploads the captured image data to the server. At this time, the image data is input and sent to the server. Once the server receives the image data, data analysis becomes possible. Specifically, this process involves data transfer over the network and confirmation of receipt.
[0820] Step 3:
[0821] The server analyzes the received image data using OCR technology. The input is image data, and the output is extracted text information. The server uses OCR software to extract character information from the image data and convert it into text data. Specifically, processes such as image preprocessing, character recognition, and text conversion are performed.
[0822] Step 4:
[0823] The server processes the user's voice and facial expression data and performs emotion analysis. Voice and facial expression data are inputs, and the output is the user's current emotional state. The server uses specialized emotion analysis software to analyze the user's stress level and changes in emotion. The specific operations include extracting voice features, recognizing facial expressions, and estimating the emotional state.
[0824] Step 5:
[0825] The server identifies areas of weakness by comparing extracted text information and emotional states with the user's learning history. The input is text information and emotional states, and the output is the identified areas of weakness. The server identifies areas where the user particularly struggles by comparing them with a database of past learning history. Specifically, this involves database matching, comparison with past performance, and reflection of emotional information.
[0826] Step 6:
[0827] The server generates similar problems and adjusts their difficulty level. The user's areas of difficulty and emotional state are inputs, and the generated problem set is the output. Emotional information is used to provide problems best suited to the user's current situation. Specific operations include problem selection, difficulty adjustment, and problem set generation.
[0828] Step 7:
[0829] The system provides explanations optimized for the user. The input is the explanation for the generated problem, and the output is the content presented to the user. The server selects the format of the explanation based on the user's emotional state and provides it in video or text format. Specifically, it handles the preparation of the explanation content, format selection, and distribution to the user.
[0830] Step 8:
[0831] Users solve similar problems. The results of their solutions become input and are sent to the server as learning progress data. The user's answers are reflected in the next learning plan as an element of progress management. Specifically, this includes answer input, result evaluation, and transmission of progress data.
[0832] Step 9:
[0833] The server records the user's learning results and creates a new learning plan. The user's answers are the input, and the newly adjusted learning plan is the output. A new learning strategy is determined based on past data and current progress. The specific actions involve updating the database, executing the plan generation algorithm, and notifying the user of the new learning strategy.
[0834] (Application Example 2)
[0835] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0836] In conventional learning systems, when a user makes a mistake, the learning experience is uniform and lacks individualization that takes into account the user's emotional state, leading to a decrease in motivation. Furthermore, providing learning materials that disregard the user's psychological state results in insufficient learning efficiency.
[0837] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0838] In this invention, the server includes an image acquisition means for taking and uploading photos of learning problems that the user has answered incorrectly, an image analysis means for analyzing the uploaded images and extracting text data, and an emotion recognition means for monitoring the user's emotional state in real time. This makes it possible to provide a personalized learning experience that corresponds to the user's emotional state and to present similar problems and explanations of appropriate difficulty levels.
[0839] "Image acquisition means" refers to a device or part of a device used to photograph learning questions that a user has answered incorrectly and to upload the image data.
[0840] "Image analysis means" refers to technologies and devices that perform processing to extract text data from uploaded image data, and which utilize optical character recognition technology.
[0841] "Analysis methods" refer to methods and technologies for identifying areas of difficulty for users by referencing their learning history based on extracted text data.
[0842] A "problem generation method" refers to a function or technology that generates similar problems based on the user's areas of weakness and adjusts the difficulty level of the problems while considering the user's emotional state.
[0843] "Explanation provision means" refers to a method or technology for preparing and providing explanations for similar problems that have been generated, in a format that is appropriate to the user's emotional state.
[0844] "Emotion recognition means" refers to devices or configurations for monitoring a user's emotional state in real time and analyzing the data.
[0845] A "feedback provision method" refers to a method or system for recording learning results and creating a new learning plan based on them.
[0846] This invention is a system for efficiently supporting user learning and aims to provide an individualized learning experience that takes into account the user's emotional state.
[0847] The system primarily consists of terminals, servers, and their respective software components. The terminals are equipped with cameras and microphones, and images are acquired when the user takes a picture of a learning problem they answered incorrectly. This image data is uploaded to a cloud server in real time, where image analysis is performed. Text data is extracted from the images using OCR technology. For optical character recognition, services such as Google Cloud Vision can be used.
[0848] Furthermore, the device has a built-in emotion recognition engine that analyzes the user's voice and facial expressions to detect their emotional state in real time. The Affectiva SDK and other tools can be used for emotion recognition. The server identifies the user's areas of difficulty based on the emotional data and uses a problem generation engine to generate similar problems tailored to the user. In this process, machine learning models and generative AI models are utilized, taking into account the user's past learning history and current emotional state.
[0849] The generated problems are accompanied by explanations in an appropriate format, taking into account the user's emotional state. For example, if the user is relaxed, a concise text explanation is provided; if they are judged to be highly focused, detailed video content is presented. This makes it possible to maximize the learning effect.
[0850] As a concrete example, consider a scenario where a user makes a mistake on an English grammar question. The user takes a picture of the question with their device, and simultaneously, an emotion recognition engine detects the user's frustration. Based on this information, the server provides relaxing video content along with simple review questions. In this way, delivering optimal learning content tailored to the user's emotions improves both motivation and efficiency in learning.
[0851] An example of a prompt might be: "Generate appropriate learning content when the user's facial expression data indicates 'frustration.' Lower the difficulty level and include relaxing video content."
[0852] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0853] Step 1:
[0854] The user takes a picture of the problem and the image is acquired. The input is the image of the problem captured by the camera, and the output is the captured image data. The terminal activates the camera at the user's command and takes a picture of the problem sheet. This image is immediately imported into the system.
[0855] Step 2:
[0856] The device uploads the captured image data to a cloud server. The input is the captured image data, and the output is the image data sent to the server. The uploaded images are ready for image analysis.
[0857] Step 3:
[0858] The server extracts text data from uploaded image data. The input is image data, and the output is extracted text data. The server uses optical character recognition (OCR) technology to recognize characters within the image and convert them into text. Through this process, the problem statement is obtained as text data.
[0859] Step 4:
[0860] The device collects and analyzes voice and facial expression data to recognize the user's emotions. The input is the user's voice and facial expression data, and the output is the analyzed emotion data. The emotion recognition engine processes this data in real time and quantifies the user's emotional state.
[0861] Step 5:
[0862] The server identifies the user's weak areas and generates problems based on text data and sentiment data. Input is text data and sentiment data, and output is similar problems and explanations tailored to the user. The optimal problem set is constructed using an AI model that analyzes past learning history and sentiment data.
[0863] Step 6:
[0864] The server optimizes and provides explanations for the generated set of problems based on the user's emotional state. The input is the user's emotional state and the generated problems, while the output is the adjusted explanatory content. The difficulty and format of the explanations are adjusted to improve the user's learning efficiency.
[0865] Step 7:
[0866] The user solves a provided similar problem and sends the results from their device to the server. The input is the user's answer, and the output is training data for feedback. The server analyzes these results and uses them to build the next training plan.
[0867] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0868] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0869] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0870] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0871] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0872] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0873] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0874] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0875] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0876] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0877] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0878] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0879] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0880] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0881] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0882] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0883] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0884] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0885] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0886] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0887] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0888] The following is further disclosed regarding the embodiments described above.
[0889] (Claim 1)
[0890] A means for users to take pictures of and upload learning problems they answered incorrectly,
[0891] An image analysis means that analyzes uploaded images and extracts text data,
[0892] An analytical method for identifying a user's weak areas from extracted text data,
[0893] A problem generation method that generates and provides similar problems based on the user's areas of weakness,
[0894] A means of providing explanations that prepare and present to users explanations for similar problems that have been generated,
[0895] A feedback system for recording learning progress and creating new learning plans,
[0896] A system that includes this.
[0897] (Claim 2)
[0898] The system according to claim 1, wherein the image analysis means has a configuration for converting a problem statement into text using optical character recognition.
[0899] (Claim 3)
[0900] The system according to claim 1, wherein the problem generation means has a configuration for adjusting the difficulty level of similar problems, taking into account the user's previous learning history.
[0901] "Example 1"
[0902] (Claim 1)
[0903] A means of acquiring information for users to take photos of and submit learning assignments they have made mistakes on,
[0904] Information analysis means for analyzing transmitted information and extracting textual information,
[0905] An evaluation method for identifying areas where a user is weak from extracted textual information,
[0906] A problem generation means that generates and provides similar problems based on the user's areas of weakness,
[0907] An explanation provision means that prepares and presents explanations for similar problems that have been generated to the user,
[0908] A means of providing an assessment that records learning progress and creates a new learning plan,
[0909] A system that includes this.
[0910] (Claim 2)
[0911] The information analysis means is configured to convert a problem statement into text using character recognition technology, according to claim 1.
[0912] (Claim 3)
[0913] The system according to claim 1, wherein the task generation means has a configuration for adjusting the difficulty level of similar tasks, taking into account the user's past learning history.
[0914] "Application Example 1"
[0915] (Claim 1)
[0916] A means of acquiring data for users to take pictures of and upload learning problems they answered incorrectly,
[0917] A data analysis method that analyzes uploaded images and extracts text data,
[0918] An information analysis method that identifies a user's weak areas from extracted text data,
[0919] Information generation means that generates and provides similar problems based on the user's areas of weakness,
[0920] Information provision means for preparing and presenting explanations for similar problems that have been generated to the user,
[0921] A means of providing results that record learning progress and create new learning programs,
[0922] A data processing means that generates easy-to-learn visual video content based on the analysis results of user weaknesses and distributes it to a data device,
[0923] A system that includes this.
[0924] (Claim 2)
[0925] The system according to claim 1, wherein the data analysis means has a configuration for converting the problem statement into text using optical character recognition.
[0926] (Claim 3)
[0927] The system according to claim 1, wherein the information generation means has a configuration for adjusting the difficulty level of similar problems, taking into account the user's past learning record.
[0928] "Example 2 of combining an emotion engine"
[0929] (Claim 1)
[0930] A means of acquiring data for users to take pictures of and upload learning problems they answered incorrectly,
[0931] A data analysis method that analyzes uploaded data and extracts text information,
[0932] An emotion analysis means for determining the user's emotional state from collected voice and facial expression data,
[0933] An analytical means for identifying a user's areas of weakness from extracted text information and emotional state,
[0934] A task generation method that generates and provides similar problems based on the user's areas of difficulty and emotional state,
[0935] A means of providing information that prepares explanations for similar problems that have been generated, and presents them in a format that is adjusted based on the user's emotional state,
[0936] A progress management tool for recording learning progress and creating new learning plans,
[0937] A system that includes this.
[0938] (Claim 2)
[0939] The data analysis means has a configuration for converting the problem into text using optical character recognition.
[0940] The system according to claim 1, wherein the emotion analysis means has a configuration for evaluating the user's emotional state from voice tone and facial expressions.
[0941] (Claim 3)
[0942] The system according to claim 1, wherein the task generation means has a configuration for adjusting the difficulty level of similar problems, taking into account the user's past learning history and emotional state.
[0943] "Application example 2 when combining with an emotional engine"
[0944] (Claim 1)
[0945] A means for users to take pictures of and upload learning problems they answered incorrectly,
[0946] An image analysis means that analyzes uploaded images and extracts text data,
[0947] An analytical method for identifying a user's weak areas from extracted text data,
[0948] A problem generation method that generates similar problems based on the user's areas of weakness and adjusts the difficulty level of the problems while considering the user's emotional state,
[0949] A means of providing explanations for generated similar problems, which are prepared and provided in a format that suits the user's emotional state.
[0950] An emotion recognition method that monitors the user's emotional state in real time,
[0951] A feedback system for recording learning progress and creating new learning plans,
[0952] A system that includes this.
[0953] (Claim 2)
[0954] The system according to claim 1, wherein the image analysis means has a configuration for converting a problem statement into text using optical character recognition.
[0955] (Claim 3)
[0956] The system according to claim 1, wherein the problem generation means has a configuration for adjusting the difficulty level of similar problems, taking into account the user's past learning history and emotional state. [Explanation of symbols]
[0957] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for users to take pictures of and upload learning problems they answered incorrectly, An image analysis means that analyzes uploaded images and extracts text data, An analytical method for identifying a user's weak areas from extracted text data, A problem generation method that generates and provides similar problems based on the user's areas of weakness, A means of providing explanations that prepare and present to users explanations for similar problems that have been generated, A feedback system for recording learning progress and creating new learning plans, A system that includes this.
2. The system according to claim 1, wherein the image analysis means has a configuration for converting the problem statement into text using optical character recognition.
3. The system according to claim 1, wherein the problem generation means has a configuration for adjusting the difficulty level of similar problems, taking into account the user's past learning history.