system

The system addresses students' reluctance to ask questions by using image and speech recognition, coupled with generative AI, to provide personalized explanations and re-explanations, enhancing their learning experience.

JP2026036282APending Publication Date: 2026-03-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024138809
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-20
Publication Date
2026-03-05

AI Technical Summary

Technical Problem

Students, particularly in junior high and high school, face difficulties in asking questions and deepening their understanding due to embarrassment and the lack of timely support, especially when studying for entrance exams.

Method used

A system that includes image recognition to analyze problem images, a generative AI to provide solutions and explanations, and speech recognition to adjust explanations based on the user's understanding, allowing users to ask questions and receive follow-up explanations at their own pace.

Benefits of technology

Enables students to ask questions and deepen their understanding by providing personalized explanations and re-explanations until comprehension is achieved, fostering a comfortable learning environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026036282000001_ABST
    Figure 2026036282000001_ABST
Patent Text Reader

Abstract

Provide a system. [Solution] means for receiving a user-provided problem image; image recognition means for analyzing a problem from the received image; means having a generating AI for processing the analyzed problem data; A means for outputting the solution and explanation obtained by the generating AI to the user; means for receiving a query from a user; a means for generating an explanation based on the received question and outputting the explanation to the user; A system including:
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In traditional educational settings, there are many cases where students find it difficult to ask questions, and where understanding does not progress with textbooks alone. This is particularly true for junior high and high school students studying for entrance exams. Even if they encounter a specific point they don't understand, they are unable to ask questions in a timely manner, and their studies continue unabated. Some students also feel embarrassed about asking questions, which hinders their learning. This issue needs to be addressed, and support is needed to help students feel comfortable asking questions and deepen their understanding. [Means for solving the problem]

[0005] The present invention solves the above-mentioned problems with a system including the following means: means for receiving an image of a problem provided by a user; image recognition means for analyzing the problem from the received image; means having a generation AI for processing the analyzed problem data; means for outputting to the user a solution and explanation obtained by the generation AI; means for receiving a question from the user; and means for generating an explanation based on the received question and outputting it to the user. Furthermore, it is desirable that the system include speech recognition means for converting speech provided by the user into text, and means for the generation AI to adjust the content of the explanation according to the user's level of understanding. This allows students to ask questions at their own pace and deepen their understanding.

[0006] ---

[0007] "User" refers to a person who uses this system to learn.

[0008] "Problem image" refers to an image of a problem that the user needs to input into the system in order to learn.

[0009] "Image recognition" refers to the technology of analyzing characters and symbols in an image and converting them into text data.

[0010] "Generative AI" refers to artificial intelligence that automatically generates solutions and explanations based on provided data.

[0011] "Solution" refers to the specific steps to solve a problem, such as mathematics.

[0012] "Explanation" refers to explanatory text that helps understand each step of the solution and the overall solution.

[0013] "Speech recognition" refers to the technology of converting voice data into text data.

[0014] "Re-explanation" refers to a new explanation provided by the generating AI based on additional questions from the user.

[0015] "Level of understanding" refers to an index showing how well the user understands the explanation.

[0016] "Convert to text" refers to the process of converting audio data into text data. [Brief explanation of the drawings]

[0017] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0018] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0019] First, the terms used in the following description will be explained.

[0020] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0021] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0022] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0023] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0025] [First embodiment]

[0026] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0027] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0028] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0029] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0030] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0032] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0033] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0034] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0035] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0036] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0037] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0038] This invention is a system that solves incomprehensible problems that arise when users study and supports deeper learning. This system provides a series of functions: the user inputs an image of a problem, and a generative AI automatically generates a solution and explanation, which are then output to the user. Furthermore, the system responds to follow-up questions from the user and provides further explanations to deepen the user's understanding.

[0039] System configuration

[0040] The system configuration is as follows:

[0041] 1. Terminal

[0042] The terminal is the device through which the user captures and inputs the image in question, and can include a smartphone, tablet, or PC.

[0043] 2. Server

[0044] The server receives and processes data sent from the device, analyzes the problem using image recognition technology, and sends the data to the generating AI.

[0045] 3. Generation AI

[0046] The AI ​​generator analyzes the received problem data and generates solutions and explanations. It also has the ability to adjust the explanations according to the user's level of understanding.

[0047] Program processing explanation

[0048] The processing of the program will be explained in natural language below.

[0049] ---

[0050] Problem capture and analysis

[0051] Terminal

[0052] The device prompts the user to take a picture of the problem, which is then saved on the device.

[0053] The device sends the saved image to the server.

[0054] server

[0055] The server receives the image sent from the terminal.

[0056] The server analyzes the received image and converts it into text data using image recognition technology (OCR).

[0057] The converted text data is formatted into a question and sent to the generation AI.

[0058] ---

[0059] Generate and display solutions and explanations

[0060] Generation AI

[0061] The generation AI analyzes the problem data it receives and generates a solution and explanation.

[0062] For example, the generated solution to the problem "2x + 3 = 7" would be the steps "2x = 7 - 3," "2x = 4," and "x = 2."

[0063] server

[0064] The server receives the solution and explanation from the generated AI.

[0065] The server sends the solution and explanation to the device.

[0066] Terminal

[0067] The terminal displays the solution and explanation to the user.

[0068] ---

[0069] Handling questions and restatements

[0070] User

[0071] If the user does not understand any part of the explanation, they can ask questions using voice recognition or text input.

[0072] Terminal

[0073] The device converts the user's speech into text.

[0074] The device sends the question to the server.

[0075] server

[0076] The server analyzes the received question and sends it to the generation AI.

[0077] Generation AI

[0078] Generative AI generates new explanations based on user questions.

[0079] server

[0080] The server receives the re-explanation from the generating AI and sends it to the terminal.

[0081] Terminal

[0082] The terminal displays the explanation again to the user.

[0083] If the user has further questions, the process is repeated.

[0084] ---

[0085] Specific examples

[0086] Example of a math problem

[0087] Consider the case where a user inputs the math problem "2x + 3 = 7." The following occurs:

[0088] 1. Problem capture

[0089] The user takes a picture of the problem using the device's camera.

[0090] The device sends the captured image to the server.

[0091] 2. Image Recognition and Problem Analysis

[0092] The server analyzes the received image and uses OCR to extract the text data "2x + 3 = 7".

[0093] The extracted text data is sent to the generation AI.

[0094] 3. Generating solutions and explanations

[0095] The generative AI generates the solutions to "2x + 3 = 7": "2x = 7 - 3," "2x = 4," and "x = 2," along with their explanations.

[0096] The server sends the generated solution and explanation to the terminal.

[0097] The terminal displays the solution and explanation to the user.

[0098] 4. Questions and Restatements

[0099] The user asks, "I don't understand why we subtract 3 from 7."

[0100] The device converts the speech into text and sends it to the server.

[0101] The server sends the question to the generation AI.

[0102] The generator AI generates a restatement: "To solve the equation, we need to shift the constant term, so we subtract 3 from 7."

[0103] The server sends the re-explanation to the terminal.

[0104] The terminal displays the re-explanation to the user.

[0105] In this way, questions can be asked and re-explained as many times as necessary until the user understands. This system allows users to learn at their own pace and achieve a deep understanding.

[0106] The processing flow will be explained below.

[0107] ---

[0108] Step 1:

[0109] The user takes a picture of the problem to be studied on the terminal.

[0110] The device will display instructions to the user saying, "Please take a picture of the problem."

[0111] The user takes a picture of the problem using the device's camera.

[0112] The device saves the captured image to its internal storage.

[0113] Step 2:

[0114] The device sends the saved image to the server.

[0115] The server receives the image data sent from the terminal.

[0116] The server temporarily stores the received image data.

[0117] Step 3:

[0118] The server uses image recognition technology (OCR) to extract text data from the image.

[0119] The server formats the extracted text data into a question format.

[0120] The server sends the formatted problem data to the generation AI.

[0121] Step 4:

[0122] The generation AI analyzes the problem data it receives.

[0123] The generative AI generates a solution to the provided problem.

[0124] Generative AI generates detailed explanations based on the solution.

[0125] Step 5:

[0126] The server receives the generated solution and explanation from the generated AI.

[0127] The server sends the received solution and explanation to the terminal.

[0128] Step 6:

[0129] The terminal displays the solution and explanation to the user.

[0130] The user checks the explanation and finds parts that he does not understand.

[0131] Step 7:

[0132] The user asks the terminal questions about anything they don't understand.

[0133] In the case of voice recognition, the user speaks a question into the terminal.

[0134] For text input: The user types the question into the device.

[0135] The device converts the user's speech into text (in the case of speech recognition).

[0136] The terminal sends a question from the user to the server.

[0137] Step 8:

[0138] The server receives the user's query.

[0139] The server sends the received question to the generation AI.

[0140] Step 9:

[0141] The generative AI generates a re-explanation based on the user's question.

[0142] The generating AI sends the new explanation to the server.

[0143] Step 10:

[0144] The server receives a re-explanation from the generated AI.

[0145] The server transmits the received explanation to the terminal.

[0146] Step 11:

[0147] The terminal displays the re-explanation to the user.

[0148] The user reviews the re-explanation and determines whether or not they understand it.

[0149] If not understood, the user asks the question again (return to step 7).

[0150] ---

[0151] By repeating this series of processes, the user can continue learning until their understanding deepens.

[0152] Example 1

[0153] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0154] In today's educational environment, users often encounter points they cannot understand while studying, and a system that can appropriately address these issues is needed. In particular, users who study independently have limited access to appropriate explanations and supplementary information. Therefore, the objective of this invention is to provide a system that can immediately resolve questions users encounter during the problem-solving process and deepen their understanding.

[0155] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0156] In this invention, the server includes a means for receiving an image of a problem provided by a user, an image recognition means for analyzing the problem from the received image, and a means having a generation AI for processing the analyzed problem data, which enables a means for a user to take a picture of the problem with a camera on a terminal and send it to the server, a means for the generation AI to analyze the problem based on the prompt and automatically generate steps for a solution and explanation, and a means for converting a user's voice input into text.

[0157] A "user" is an entity that uses the system to enter problems and receive solutions and explanations.

[0158] A "problem image" is a digital image that visually represents the problem the user wants to solve.

[0159] A "terminal" is a device that a user uses to capture and input the image in question, and includes smartphones, tablets, personal computers, etc.

[0160] A "server" is the central part of the system that receives and processes data sent from the terminals.

[0161] "Image recognition means" refers to techniques and devices that extract text data from received images.

[0162] "Text data" is character information extracted by image recognition means.

[0163] "Generative AI" is artificial intelligence that generates solutions and explanations based on analyzed problem data.

[0164] A "solution" is a procedure and computational process for solving a given problem.

[0165] "Explanation" is a text that explains each step of the solution and the background and reasons behind it.

[0166] A "prompt" is a form of text or data that is given to a generative AI as instructions or input to solve a problem.

[0167] "Voice input" refers to the act of inputting voice uttered by a user as information.

[0168] "Speech recognition means" refers to the technology and devices that convert a user's voice input into text.

[0169] A "re-explanation" is a new explanation generated by the AI ​​based on additional questions from the user.

[0170] This invention is a system that solves problems that arise when users are studying and supports deepening their learning. This system provides a series of functions in which the user inputs an image of a problem, and a generative AI automatically generates a solution and explanation, which are then output to the user. In addition, the system has a mechanism that responds to follow-up questions from the user and provides further explanations to deepen the user's understanding.

[0171] System configuration

[0172] 1. Terminal

[0173] The terminal is a device that users use to take and input images of the problem, and mainly includes smartphones, tablets, and PCs. It temporarily stores the images of the problem taken by the user and then sends them to the server.

[0174] 2. Server

[0175] The server receives the image sent from the device and uses Optical Character Recognition (OCR) technology to analyze the received image. For example, it uses a library such as Tesseract OCR to extract text data from the image. The extracted text data is then sent to the generation AI.

[0176] 3. Generation AI

[0177] The generating AI analyzes the problem according to the prompts based on the text data sent from the server, and generates a solution and explanation. The generating AI can use a generative model such as GPT. The generated solution and explanation are sent to the device via the server.

[0178] 4. Voice Recognition Methods

[0179] When a user has questions about something they don't understand, they can ask using voice recognition or text input. The device converts the user's voice into text. Voice recognition can use services such as Google® Speech-to-Text API or Amazon Transcribe.

[0180] Specific examples

[0181] Consider the case where a user inputs a math problem, "2x + 3 = 7." The system operates as follows:

[0182] 1. Problem capture

[0183] The user takes the image in question using the device's camera. At this time, the device launches the smartphone's camera app and displays a "Take a Photo" button. When the user presses the "Take a Photo" button, the image is captured and temporarily saved in the device's memory. The device then sends the image data to the server.

[0184] 2. Image Recognition and Problem Analysis

[0185] The server passes the received image data to a local image analysis engine, which uses OCR technology to extract the text data "2x + 3 = 7" from the image. This text data is then sent to the generation AI.

[0186] 3. Generating solutions and explanations

[0187] The generative AI analyzes the problem based on the prompt and generates the solutions "2x = 7 - 3", "2x = 4", "x = 2" and detailed step-by-step explanations. For example, the following prompts are used:

[0188] Please solve the equation 2x + 3 = 7 and explain each step.

[0189] The generated solution and explanation are sent to the terminal via the server.

[0190] 4. Displaying the solution and explanation

[0191] The device then displays the received solutions and explanations in a user-friendly format on the screen, with each step displayed step-by-step and accompanied by a detailed explanation for that step.

[0192] 5. Questions and Restatements

[0193] The user asks, "I don't understand why 3 is subtracted from 7." The device converts the user's voice into text using speech recognition technology and sends it to the server. The server sends the question as a prompt to the generation AI, which then generates a new explanation. This new explanation is sent to the device via the server and displayed to the user again.

[0194] In this way, questions and re-explanations can be repeated as many times as necessary until the user understands. Through this system, users can learn at their own pace and achieve a deep understanding.

[0195] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0196] Step 1: Take a picture of the problem and send it to us

[0197] Terminal

[0198] The user takes the image in question using the device's camera. For example, they launch the camera app on their smartphone and a "Take Photo" button appears.

[0199] When the user presses the "take a photo" button, an image is captured and temporarily saved in the device's internal memory.

[0200] The device sends the stored image data to the server using an HTTP POST request.

[0201] Input: A user-taken image of the problem

[0202] Output: Image data sent to the server

[0203] Specific behavior:

[0204] 1. The user activates the device camera.

[0205] 2. Press the "Capture" button to capture the image in question.

[0206] 3. Temporarily save image data.

[0207] 4. Send the image data to the server.

[0208] Step 2: Receiving and analyzing images

[0209] server

[0210] The server saves the image data received via the HTTP request in local storage.

[0211] The server passes the saved image data to an OCR (Optical Character Recognition) engine and begins processing.

[0212] OCR technology (e.g., Tesseract OCR) is used to extract text data from the image.

[0213] The extracted text data is formatted into questions and prompts are created to be sent to the generation AI.

[0214] Input: Image data sent from the device

[0215] Output: Problem text data to send to the generation AI

[0216] Specific behavior:

[0217] 1. The server receives the image data.

[0218] 2. Save the received image data locally.

[0219] 3. Pass the image data to the OCR engine.

[0220] 4. Extract text data using OCR technology.

[0221] 5. Format the text data into a question format.

[0222] 6. Create a prompt to send the formatted text data to the generation AI.

[0223] Step 3: Generate solutions and explanations

[0224] Generation AI

[0225] The generation AI receives the text data sent from the server.

[0226] The system analyzes the received text data as prompts and automatically generates steps to solve the problem and detailed explanations for each.

[0227] The generated solutions are in the form of, for example, "2x = 7 - 3," "2x = 4," or "x = 2." The explanation is something like, "First, subtract 3 on the left side from 7 on the right side."

[0228] Input: Question text data sent from the server

[0229] Output: Generated solution and explanation

[0230] Specific behavior:

[0231] 1. The generation AI receives the text data.

[0232] 2. Analyze text data as prompts.

[0233] 3. Automatically generate solutions and detailed explanations.

[0234] Step 4: Submit and view your solution and explanation

[0235] server

[0236] The server receives the generated solution and explanation from the generating AI.

[0237] The integrity of the received data is checked and it is sent to the terminal in JSON format or similar.

[0238] Input: Solution and explanation sent from the generating AI

[0239] Output: Solution and explanation data sent to the terminal

[0240] Specific behavior:

[0241] 1. The server receives the solution and explanation.

[0242] 2. Check data integrity.

[0243] 3. The data whose integrity has been confirmed is sent to the terminal.

[0244] Terminal

[0245] The terminal analyzes the solutions and explanations received from the server and displays them on the screen in a format that is easy for the user to see.

[0246] Each step is displayed step by step and a detailed explanation for each step is also provided.

[0247] Input: Solution and explanation data sent from the server

[0248] Output: The solution and explanation displayed to the user

[0249] Specific behavior:

[0250] 1. The device receives the solution and explanation.

[0251] 2. The solution and explanation are analyzed and displayed on the screen.

[0252] Step 5: Submit your question and voice input

[0253] User

[0254] If the user does not understand the explanation, they can enter a question using voice recognition or text input.

[0255] For example, tap the microphone icon on your smartphone to start voice input and say, "I don't know why I subtract 3 from 7."

[0256] Input: Additional questions about the explanation

[0257] Output: Input question data to the terminal

[0258] Specific behavior:

[0259] 1. The user taps the microphone icon.

[0260] 2. The user types in a question by voice.

[0261] Terminal

[0262] The device converts the user's speech into text using services such as the Google Speech-to-Text API or Amazon Transcribe.

[0263] The converted text data is sent to the server.

[0264] Input: A spoken question from the user

[0265] Output: Sends text-formatted question data to the server

[0266] Specific behavior:

[0267] 1. The device receives the audio data.

[0268] 2. Convert voice to text using voice recognition technology.

[0269] 3. Send the text data to the server.

[0270] Step 6: Parse the question and generate a restatement

[0271] server

[0272] The server receives the question sent from the terminal.

[0273] The received text data is sent to the generation AI as a prompt.

[0274] Input: Text question data sent from the terminal

[0275] Output: Prompt data to send to the generation AI

[0276] Specific behavior:

[0277] 1. The server receives the question text data.

[0278] 2. Send the text data to the generation AI as a prompt.

[0279] Generation AI

[0280] Generative AI generates new explanations based on user questions.

[0281] For example, it generates a restatement like "To solve the equation, we need to shift the constant term, so subtract 3 from 7."

[0282] Input: Question prompt sent from the server

[0283] Output: Generated restatement

[0284] Specific behavior:

[0285] 1. The generative AI receives a question prompt.

[0286] 2. Automatically generate new explanations based on questions.

[0287] Step 7: Submit and view the recap

[0288] server

[0289] The server receives the generated re-explanation from the generation AI and sends it to the terminal.

[0290] Input: Restatement sent by the generating AI

[0291] Output: Recap what is sent to the terminal

[0292] Specific behavior:

[0293] 1. The server receives the re-explanation.

[0294] 2. Send the re-explanation to the device.

[0295] Terminal

[0296] The terminal displays the re-explanation received from the server to the user.

[0297] Input: Recap sent by server

[0298] Output: A restatement that is displayed to the user

[0299] Specific behavior:

[0300] 1. The device receives the re-explanation.

[0301] 2. Display the re-explanation to the user.

[0302] In this way, the user can ask questions and receive re-explanations as many times as necessary until they understand. This system allows users to learn at their own pace and achieve a deep understanding.

[0303] (Application example 1)

[0304] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0305] In factories, work procedures are complex, and employees often make mistakes, leading to a decline in production efficiency and the inability to quickly resolve problems that arise during the work process. Newly introduced manuals and work procedures also tend to be poorly understood. A system is needed that allows employees to easily understand work procedures and receive assistance in resolving problems.

[0306] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0307] In this invention, the server includes means for receiving an image of a task provided by a user, image recognition means for analyzing the task from the received image, means having a generation AI for processing the analyzed task data, means for outputting to the user a solution and explanation obtained by the generation AI, means for receiving questions from the user, means for generating an explanation based on the received question and outputting it to the user, and means for displaying the generated solution and procedure through a smart device to support complex work procedures and problem solving in a factory. This enables employees to quickly solve problems they encounter during work and work efficiently.

[0308] The "means for receiving images of tasks provided by users" includes a device that allows a factory employee to take an image of a work procedure manual, a drawing, etc., and transmit the image to a server.

[0309] "Image recognition means" refers to technology that analyzes received images, identifies and extracts text and figures from the images, and converts them into data.

[0310] "Means having generative AI" refers to an artificial intelligence (AI) system that analyzes input data and automatically generates solutions and explanations.

[0311] "Means for outputting solutions and explanations to users" refers to technology for providing the solutions and explanations generated by the generation AI to factory employees via display devices or audio output devices.

[0312] The term "means for receiving questions" refers to technology that provides an interface for factory employees to input or voice questions and receives those questions.

[0313] "Means of generating new explanations based on questions and outputting them to users" refers to technology that uses a generative AI to reanalyze received questions, generate new explanations and procedures, and provide them to factory employees.

[0314] "Means for displaying generated solutions and procedures through smart devices to support complex work procedures and problem-solving within factories" refers to technology that uses wearable devices such as smart glasses or tablet terminals to display generated procedures and solutions to factory employees in real time.

[0315] The present invention is a system for supporting complex work procedures and problem solving within a factory, and in particular provides work procedures and solutions to factory employees using wearable devices such as smart glasses and tablet terminals.

[0316] System configuration

[0317] The system configuration is as follows:

[0318] 1. Terminal

[0319] Terminals are devices that factory employees use to take and input images of work procedures and drawings. These devices include smart glasses and tablets. These terminals are equipped with a camera function and send the images to a server.

[0320] 2. Server

[0321] The server is responsible for receiving and processing data sent from the terminal. The server has the following functions:

[0322] Image recognition means: The received image is analyzed and converted into text data using OCR (optical character recognition) technology. This converted data is sent to the generation AI.

[0323] Question analysis means: Receives and analyzes questions from factory employees. This data is also sent to the generation AI.

[0324] 3. Generation AI

[0325] The generative AI analyzes the received problem data and questions, and generates solutions and explanations. It also has the ability to adjust the explanation content according to the user's level of understanding. Examples of generative AI models used include GPT-4 (registered trademark).

[0326] Program processing explanation

[0327] Problem capture and analysis

[0328] 1. The device instructs factory employees to take pictures of work procedures and drawings, which are then saved on the device.

[0329] 2. The device sends the saved image to the server.

[0330] 3. The server receives the image sent from the device.

[0331] 4. The server analyzes the received image and converts it into text data using OCR.

[0332] 5. The converted text data is formatted into a question and sent to the generation AI.

[0333] Generate and display solutions and explanations

[0334] 1. The generation AI analyzes the problem data it receives and generates specific work procedures and solutions.

[0335] 2. The generated procedures and solutions are sent to the terminal via the server.

[0336] 3. The terminal displays the generated solution and procedures to the factory worker.

[0337] Handling questions and restatements

[0338] 1. If a factory employee has questions about parts of the explanation that they don't understand, they can ask using voice recognition or text input.

[0339] 2. The device converts the user's speech into text.

[0340] 3. The device sends the question to the server.

[0341] 4. The server analyzes the received question and sends it to the generation AI.

[0342] 5. Generative AI generates new explanations based on the user's questions.

[0343] 6. The server receives the re-explanation from the generation AI and sends it to the device.

[0344] 7. The device displays the explanation again to the user.

[0345] Specific examples

[0346] Work procedure support example

[0347] Consider a case where a factory worker does not understand the work procedure "attach part A to part B." The following process is performed:

[0348] 1. A factory employee uses the device's camera to take an image of a work procedure manual or drawing.

[0349] 2. The device sends the captured image to the server.

[0350] 3. The server analyzes the received image and extracts text data using OCR.

[0351] 4. Send the extracted text data to the generation AI.

[0352] 5. The generative AI generates a solution and explanation such as "Steps for assembling parts A and B: 1. Remove the screws from part A. 2. Attach part A to part B."

[0353] 6. The generated procedure is sent to the terminal via the server and displayed to the factory employee.

[0354] 7. A factory worker asks, "Which way do you turn the screw?"

[0355] 8. The question is converted into text using voice recognition technology and sent to the server.

[0356] 9. The server sends the question to the generation AI, which generates a re-explanation: "Rotate clockwise."

[0357] 10. The re-explanation is sent to the terminal and displayed to the factory employee.

[0358] As a result, the system of the present invention can improve the work efficiency of users in a factory and speed up the problem solving process.

[0359] Prompt Sentence Examples

[0360] "Please solve the following problem: I don't understand step C: "Step C: Attach part A to part B.""

[0361] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0362] Step 1:

[0363] The terminal instructs the factory employee to take an image of the work procedure manual or blueprint. The factory employee takes an image of the work procedure manual or blueprint using the camera on the terminal. The input of this step is the factory employee's photographing behavior and the captured image, and the output is an image file saved on the terminal.

[0364] Step 2:

[0365] The device sends the saved image to the server. The device sends the image file taken by the factory employee to the server using an HTTP POST request. The input of this step is the saved image file, and the output is the image data received by the server.

[0366] Step 3:

[0367] The server receives the image sent from the terminal. The server temporarily stores the received image data. The input of this step is the image data sent from the terminal, and the output is the image data stored on the server.

[0368] Step 4:

[0369] The server analyzes the received image and converts it into text data using OCR. The server uses OCR technology (e.g., pytesseract) to extract text information from the image. The input for this step is the image data stored on the server, and the output is the extracted text data.

[0370] Step 5:

[0371] The server formats the converted text data into a question format and sends it to the generation AI. The server formats the text data into an appropriate format (e.g., JSON format) and sends an API request to the generation AI. The input of this step is the extracted text data, and the output is the question format data sent to the generation AI.

[0372] Step 6:

[0373] The generative AI analyzes the received problem data and generates a specific work procedure or solution. The generative AI (e.g., GPT-4) generates a solution and explanation based on the problem data. The input of this step is data in the form of a problem, and the output is the generated solution and explanation.

[0374] Step 7:

[0375] The server sends the generated solution and explanation to the terminal. The server sends the data received from the generation AI to the terminal. The input of this step is the generated solution and explanation, and the output is the solution and explanation sent to the terminal.

[0376] Step 8:

[0377] The terminal displays the generated solution and explanation to the factory employee. The terminal displays the received solution and explanation on the display. The input of this step is the solution and explanation sent to the terminal, and the output is display information that can be viewed by the factory employee.

[0378] Step 9:

[0379] If a factory employee does not understand a part of the explanation, they ask a question using voice recognition or text input. The factory employee then uses the terminal's voice input function or keyboard to input the question. The input for this step is the factory employee's voice or text input, and the output is the question data recognized by the terminal.

[0380] Step 10:

[0381] The terminal converts the user's voice into text. The terminal converts the voice data into text using voice recognition technology. The input of this step is the voice data of the factory employee, and the output is text data.

[0382] Step 11:

[0383] The terminal sends the question content to the server. The terminal then sends the question content converted into text data to the server. The input to this step is the text data, and the output is the question data sent to the server.

[0384] Step 12:

[0385] The server analyzes the received question and sends it to the generation AI. The server then formats the question and sends an API request to the generation AI again. The input of this step is the question data received by the server, and the output is the question data sent to the generation AI.

[0386] Step 13:

[0387] The generation AI generates a new explanation based on the user's question. The generation AI analyzes the question data and generates an appropriate re-explanation. The input of this step is the question data sent to the generation AI, and the output is the generated re-explanation data.

[0388] Step 14:

[0389] The server receives the re-explanation from the generation AI and sends it to the terminal. The server sends the generated re-explanation data to the terminal. The input of this step is the generated re-explanation data, and the output is the re-explanation data sent to the terminal.

[0390] Step 15:

[0391] The terminal displays the re-explanation to the factory employee. The terminal displays the re-explanation data on a display and provides it to the factory employee. The input of this step is the re-explanation data sent to the terminal, and the output is re-explanation information that can be viewed by the factory employee.

[0392] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0393] This invention is an educational system that not only solves problems that users do not understand when studying, but also recognizes the user's emotions and provides appropriate instruction methods. This system provides a series of functions: the user inputs an image of a problem, and a generative AI generates a solution and explanation, which is then output to the user. Furthermore, by incorporating a new emotion engine, the system has the function of recognizing the user's emotions and adjusting the explanation content according to the user's emotions.

[0394] System configuration

[0395] The system configuration is as follows:

[0396] 1. Terminal

[0397] The terminal is the device through which the user captures and inputs the image in question, and can include smartphones, tablets, and personal computers.

[0398] 2. Server

[0399] The server receives and processes data sent from the device, analyzes the problem using image recognition technology, and sends the data to the generating AI.

[0400] 3. Generation AI

[0401] The AI ​​generator analyzes the received problem data and generates solutions and explanations. It also has the ability to adjust the explanations according to the user's level of understanding.

[0402] 4. Emotion Engine

[0403] The emotion engine recognizes the user's emotions and provides a means for the system to respond based on those emotions. Specifically, it analyzes voice and facial expressions to determine whether the user understands or is confused.

[0404] Program processing explanation

[0405] The processing of the program will be explained in natural language below.

[0406] ---

[0407] Problem capture and analysis

[0408] Terminal

[0409] The device prompts the user to take a picture of the problem, which is then saved on the device.

[0410] The user takes a picture of the problem using the device's camera function.

[0411] The device sends the saved image to the server.

[0412] server

[0413] The server receives the image data sent from the terminal.

[0414] The server temporarily stores the received image data.

[0415] The server uses image recognition technology (OCR) to extract text data from the image.

[0416] The extracted text data is formatted into a question and sent to the generation AI.

[0417] ---

[0418] Generate and display solutions and explanations

[0419] Generation AI

[0420] The generation AI analyzes the problem data it receives.

[0421] The generative AI generates a solution to the provided problem.

[0422] Generative AI generates detailed explanations based on the solution.

[0423] server

[0424] The server receives the generated solution and explanation from the generated AI.

[0425] The server sends the received solution and explanation to the terminal.

[0426] Terminal

[0427] The terminal displays the solution and explanation to the user.

[0428] The user checks the explanation and finds parts that he does not understand.

[0429] ---

[0430] Handling questions and restatements

[0431] User

[0432] The user asks the terminal questions about anything they don't understand.

[0433] In the case of voice recognition, the user speaks a question into the terminal.

[0434] For text input: The user types the question into the device.

[0435] Terminal

[0436] The device converts the user's speech into text (in the case of speech recognition).

[0437] The terminal sends a question from the user to the server.

[0438] server

[0439] The server receives the user's query.

[0440] The server sends the received question to the generation AI.

[0441] Generation AI

[0442] The generative AI generates a re-explanation based on the user's question.

[0443] The generating AI sends the new explanation to the server.

[0444] server

[0445] The server receives a re-explanation from the generated AI.

[0446] The server transmits the received explanation to the terminal.

[0447] Terminal

[0448] The terminal displays the re-explanation to the user.

[0449] The user reviews the re-explanation and determines whether or not they understand it.

[0450] If the user does not understand, he or she asks the question again.

[0451] ---

[0452] Recognizing and Responding to Emotions

[0453] Terminal

[0454] The device collects the user's voice and camera footage.

[0455] The terminal transmits the collected data to the server.

[0456] server

[0457] The server sends the received data to the emotion engine.

[0458] The emotion engine analyzes voice and facial expressions to recognize the user's emotions.

[0459] The emotion engine sends the recognition results to the server.

[0460] Generation AI

[0461] The generative AI adjusts the commentary content based on the recognition results.

[0462] For example, if the user is confused, generate a more detailed explanation.

[0463] server

[0464] The server receives the adjusted commentary from the generated AI.

[0465] The server sends the adjusted commentary to the terminal.

[0466] Terminal

[0467] The terminal displays the adjusted description to the user.

[0468] ---

[0469] Specific examples

[0470] Example of a math problem

[0471] Consider the case where a user inputs the math problem "2x + 3 = 7." The following occurs:

[0472] 1. Problem capture

[0473] The user takes a picture of the problem using the device's camera.

[0474] The device sends the captured image to the server.

[0475] 2. Image Recognition and Problem Analysis

[0476] The server analyzes the received image and uses OCR to extract the text data "2x + 3 = 7".

[0477] The extracted text data is sent to the generation AI.

[0478] 3. Generating solutions and explanations

[0479] The generative AI generates the solutions to "2x + 3 = 7": "2x = 7 - 3," "2x = 4," and "x = 2," along with their explanations.

[0480] The server sends the generated solution and explanation to the terminal.

[0481] The terminal displays the solution and explanation to the user.

[0482] 4. Questions and Restatements

[0483] The user asks, "I don't understand why we subtract 3 from 7."

[0484] The device converts the speech into text and sends it to the server.

[0485] The server sends the question to the generation AI.

[0486] The generator AI generates a restatement: "To solve the equation, we need to shift the constant term, so we subtract 3 from 7."

[0487] The server sends the re-explanation to the terminal.

[0488] The terminal displays the re-explanation to the user.

[0489] 5. Emotion Recognition and Response

[0490] The device collects the user's voice and camera footage and sends them to the server.

[0491] The emotion engine analyzes voice and facial expressions to recognize the user's emotions.

[0492] The generative AI adjusts the commentary based on the perceived emotion (generating a more detailed explanation if the person is confused).

[0493] The server sends the adjusted commentary to the terminal, which displays it to the user.

[0494] In this way, users can receive support tailored to their level of understanding and emotional state. This system allows users to learn at their own pace and achieve deep understanding.

[0495] The processing flow will be explained below.

[0496] I understand. Now, I will explain the process flow of the invention that combines the emotion engine in concrete steps.

[0497] ---

[0498] Step 1:

[0499] The device will prompt the user to "Take a picture of the problem."

[0500] The user uses the device's camera to take the image in question.

[0501] The device saves the captured image to its internal storage.

[0502] Step 2:

[0503] The device sends the saved image to the server.

[0504] The server receives the image data sent from the terminal.

[0505] The server temporarily stores the received image data.

[0506] Step 3:

[0507] The server uses image recognition technology (OCR) to extract text data from the image.

[0508] The server formats the extracted text data into a question format.

[0509] The server sends the formatted problem data to the generation AI.

[0510] Step 4:

[0511] The generation AI analyzes the problem data it receives.

[0512] The generative AI generates a solution to the provided problem.

[0513] Generative AI generates detailed explanations based on the solution.

[0514] Step 5:

[0515] The server receives the generated solution and explanation from the generated AI.

[0516] The server sends the received solution and explanation to the terminal.

[0517] Step 6:

[0518] The terminal displays the solution and explanation to the user.

[0519] The user checks the explanation and finds parts that he does not understand.

[0520] Step 7:

[0521] The user asks the terminal questions about anything they don't understand.

[0522] In the case of voice recognition, the user speaks a question into the terminal.

[0523] For text input: The user types the question into the device.

[0524] Step 8:

[0525] The device converts the user's speech into text (in the case of speech recognition).

[0526] The terminal sends a question from the user to the server.

[0527] Step 9:

[0528] The server receives the user's query.

[0529] The server sends the received question to the generation AI.

[0530] Step 10:

[0531] The generative AI generates a re-explanation based on the user's question.

[0532] The generating AI sends the new explanation to the server.

[0533] Step 11:

[0534] The server receives a re-explanation from the generated AI.

[0535] The server transmits the received explanation to the terminal.

[0536] Step 12:

[0537] The terminal displays the re-explanation to the user.

[0538] The user reviews the re-explanation and determines whether or not they understand it.

[0539] If not understood, the user asks the question again (return to step 7).

[0540] Step 13:

[0541] The device collects the user's voice and camera footage.

[0542] The terminal transmits the collected data to the server.

[0543] Step 14:

[0544] The server sends the received data to the emotion engine.

[0545] The emotion engine analyzes voice and facial expressions to recognize the user's emotions.

[0546] The emotion engine sends the recognition results to the server.

[0547] Step 15:

[0548] The generative AI adjusts the commentary content based on the recognition results.

[0549] For example, if the user is confused, generate a more detailed explanation.

[0550] Step 16:

[0551] The server receives the adjusted commentary from the generated AI.

[0552] The server sends the adjusted commentary to the terminal.

[0553] Step 17:

[0554] The terminal displays the adjusted description to the user.

[0555] ---

[0556] Through this series of processes, users can receive support tailored to their level of understanding and emotional state.

[0557] Example 2

[0558] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0559] In recent years, there has been a growing need for educational support systems that can quickly and accurately resolve user issues. However, conventional systems lack the explanations and suggestions needed to resolve issues that users cannot understand. Furthermore, because they do not take the user's emotional state into consideration, learning effectiveness may not be fully realized. To solve these problems, there is a need for the development of a system that not only provides answers to users' questions but also recognizes the user's emotions and provides appropriate guidance accordingly.

[0560] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0561] In this invention, the server includes means for receiving images of problems provided by the user, image recognition means for analyzing the problems from the received images, artificial intelligence means for processing the analyzed problem data, means for outputting to the user the solutions and explanations obtained by the artificial intelligence means, means for receiving questions from the user, means for generating new explanations based on the received questions and outputting them to the user, emotion recognition means for collecting and analyzing emotion data of the user, and means for adjusting the content of the explanations based on the analysis results. This makes it possible to solve the user's understanding problems and provide appropriate guidance according to the user's emotional state, thereby maximizing the learning effect.

[0562] "User" refers to a person who uses the system.

[0563] "Question image" refers to an image file containing a study question that the user has taken using the camera function.

[0564] "Means for receiving" refers to the function that allows the system to receive problem images and questions provided by the user.

[0565] "Image recognition means" refers to technology (e.g., OCR technology) for extracting text data or problem information from received images.

[0566] "Generative AI" refers to a program that generates solutions and explanations from problem data analyzed using artificial intelligence technology.

[0567] "Means of output" refers to functions for providing users with the solutions and explanations obtained by the generating AI, such as screen display and audio output functions.

[0568] "Means for receiving questions" refers to a function for receiving additional questions or inquiries from users.

[0569] "Means for generating a new explanation" refers to the function of the generation AI to create a new explanation in response to a question from the user.

[0570] "Emotion data" refers to information about emotions analyzed from the user's voice, video, etc.

[0571] "Emotion recognition means" refers to technology for analyzing collected emotion data and recognizing the user's emotional state.

[0572] The "adjustment means" refers to a function for appropriately changing the commentary content based on the analysis results of the emotion recognition means.

[0573] This invention is an educational support system that not only resolves points that a user does not understand when studying, but also recognizes the user's emotions and provides appropriate teaching methods. This system is a combination of various hardware and software and consists of the following components:

[0574] First, the devices used by users include devices such as smartphones, tablets, and PCs. Users can use these devices to take pictures of study questions. The devices use their camera functions to acquire the images of the questions taken by the users.

[0575] The acquired image data is then sent to a server via a network. The server stores the received image data and uses image recognition technology (OCR = Optical Character Recognition) to extract text data from the image. This text data is then analyzed to determine the content of the problem.

[0576] The extracted text data is sent to the generative AI, which analyzes the problem based on the text data and generates a solution and a detailed explanation. The generative AI uses the latest machine learning and natural language processing technologies, enabling advanced analysis and generation.

[0577] The generated solutions and explanations are then sent to the terminal via the server. The terminal displays the received solutions and explanations to the user. The user can check the solutions and explanations displayed and ask additional questions about any parts they do not understand.

[0578] A user's question is entered into the device through text input or voice input. The device that receives the question sends the data to the server. The server then sends the user's question to the generation AI, which then generates a re-explanation. The generated re-explanation is sent to the device via the server and displayed to the user.

[0579] The system also incorporates emotion recognition functionality. The device collects the user's voice and camera footage and sends it to a server. The server then uses an emotion engine to analyze the user's emotion data and identify their emotional state. The emotion recognition data is then sent to a generation AI, which then adjusts the commentary content according to the user's emotional state. For example, if the user is confused, a more detailed explanation will be generated.

[0580] As a concrete example, consider the case where a user inputs the math problem "2x + 3 = 7." The user takes an image of the problem using the device's camera and sends it from the device to the server. The server uses OCR technology to extract the text data "2x + 3 = 7" and sends it to the generation AI. The generation AI generates a solution and detailed explanation for "2x + 3 = 7" and sends it to the device via the server. The device then displays the solution and explanation to the user.

[0581] If a user asks, "I don't know why we subtract 3 from 7," the device sends the question to the server, and the generation AI generates a restatement, saying, "Because we need to shift the constant term to solve the equation." The server sends the restatement to the device, which displays it to the user.

[0582] Here are some examples of specific prompts:

[0583] Problem: Generate a solution and explanation for "2x + 3 = 7".

[0584] User Question: "I don't understand why we subtract 3 from 7."

[0585] User sentiment: Confused

[0586] Please provide a more detailed explanation.

[0587] This system allows users to receive learning support at their own pace, and by receiving guidance tailored to their emotional state, they can achieve a deeper understanding.

[0588] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0589] Step 1:

[0590] The device instructs the user to take a picture of the problem. Specifically, the device displays "Please take a picture of the problem." Input: User's instruction appears. Output: Message is displayed.

[0591] Step 2:

[0592] The user takes a picture of the problem using the device's camera. Input: The image of the problem taken by the user. Output: Image data.

[0593] Step 3:

[0594] The image captured by the device is sent to the server. Input: Image data from the device. Output: Image file sent to the server.

[0595] Step 4:

[0596] The server receives the image data. Input: Image data sent from the terminal. Output: Reception confirmation signal and storage of the image data.

[0597] Step 5:

[0598] The server uses image recognition technology (OCR) to extract text data. The server launches image analysis software to extract text data from the image. Input: Received image data. Output: Extracted text data.

[0599] Step 6:

[0600] The server sends the extracted text data to the generation AI. The server calls the generation AI's API and sends the text data. Input: Text data. Output: Data sent to the generation AI.

[0601] Step 7:

[0602] The generative AI generates a solution and explanation for the problem. The generative AI analyzes the input text data and generates an appropriate solution and detailed explanation. Input: Submitted text data. Output: Generated solution and explanation data.

[0603] Step 8:

[0604] The server sends the generated solution and explanation to the terminal. The server transfers the solution and explanation data received from the generating AI to the terminal. Input: Solution and explanation data from the generating AI. Output: Sending the solution and explanation data to the terminal.

[0605] Step 9:

[0606] The terminal displays the solution and explanation to the user. The solution and explanation are displayed on the terminal screen. Input: Solution and explanation data. Output: Display to user.

[0607] Step 10:

[0608] The user inputs a question. The user checks the solution and explanation and inputs a question about the part they don't understand. Voice input or text input is available. Input: User's question (text or voice). Output: Question data.

[0609] Step 11:

[0610] The terminal sends a question to the server. Input: Question data. Output: Question data sent to the server.

[0611] Step 12:

[0612] The server sends the question to the generation AI. The server calls the generation AI's API and sends the question data. Input: User's question data. Output: Sending question data to the generation AI.

[0613] Step 13:

[0614] The generation AI generates a restatement based on the question. The generation AI analyzes the question data and generates a restatement for the question. Input: Submitted question data. Output: Generated restatement data.

[0615] Step 14:

[0616] The server sends the re-explanation to the terminal. The server transfers the re-explanation data received from the generating AI to the terminal. Input: Re-explanation data from the generating AI. Output: Sending the re-explanation data to the terminal.

[0617] Step 15:

[0618] The terminal displays the re-explanation to the user. The re-explanation is displayed on the terminal screen. Input: Re-explanation data. Output: Re-display to user.

[0619] Step 16:

[0620] The device collects the user's emotional data and sends it to the server. The device collects the user's audio and video data and sends it to the server. Input: User's audio and video data. Output: Sending emotional data to the server.

[0621] Step 17:

[0622] The server sends emotion data to the emotion engine. Input: Emotion data sent from the device. Output: Data sent to the emotion engine.

[0623] Step 18:

[0624] The emotion engine analyzes the user's emotions. The emotion engine analyzes voice and facial expressions to recognize the user's emotional state. Input: Emotion data. Output: Analyzed emotional state data.

[0625] Step 19:

[0626] The emotion engine sends the analysis results to the generation AI. The emotion engine sends the analysis results to the generation AI, which adjusts the commentary content. Input: Analysis result data. Output: Data sent to the generation AI.

[0627] Step 20:

[0628] The generation AI adjusts the commentary content according to the emotion. The generation AI adjusts the commentary content based on the recognition results and generates an appropriate commentary. Input: Analysis result data. Output: Adjusted commentary data.

[0629] Step 21:

[0630] The server sends the adjusted commentary to the terminal. The server transfers the adjusted commentary data received from the generation AI to the terminal. Input: Adjusted commentary data from the generation AI. Output: Sending adjusted commentary data to the terminal.

[0631] Step 22:

[0632] The terminal displays the adjusted commentary to the user. The adjusted commentary is displayed on the terminal screen. Input: Adjusted commentary data. Output: Adjusted commentary displayed to the user.

[0633] (Application example 2)

[0634] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0635] When a user encounters a point in the educational process that they do not understand, a common educational system has difficulty quickly recognizing that confusion and responding effectively. In particular, flexible responses based on the learner's emotions and level of understanding are required, but this has been difficult to achieve with conventional systems. In addition, there is a lack of means to grasp the learner's progress in real time and provide appropriate feedback on the spot.

[0636] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0637] In this invention, the server includes means for receiving an image of a problem provided by a user, image recognition means for analyzing the problem from the received image, means having a generation AI for processing the analyzed problem data, means for outputting to the user a solution and explanation obtained by the generation AI, means for receiving a question from the user, means for generating an explanation based on the received question and outputting it to the user, means for identifying the user's emotion, and means for adjusting the content of the explanation based on the identified emotion. This enables flexible learning support according to the user's emotion and level of understanding, and also enables real-time progress confirmation and feedback.

[0638] "User" refers to the subject who uses this system to study, and who performs operations such as entering questions and asking questions.

[0639] "Means for receiving problem images" refers to an interface for receiving problem images photographed or scanned by a user and importing them into the system.

[0640] "Image recognition means" refers to technology that analyzes the received image of a question and extracts the question content, such as text data.

[0641] "Generative AI" refers to artificial intelligence technology that generates solutions and explanations based on extracted problem data.

[0642] "Means for outputting to the user" refers to an interface for visually or audibly conveying the solution and explanation obtained by the generative AI to the user.

[0643] The "means for receiving a question" refers to an interface for receiving a question within the system when a user inputs a question about a point that he or she does not understand.

[0644] The "means for generating a new explanation and outputting it to the user" refers to an interface for generating a new explanation based on a received user question and conveying that explanation to the user visually or audibly.

[0645] "Means for identifying user emotions" refers to technology that analyzes the user's facial expressions and voice to identify the emotion the user is currently feeling.

[0646] The "means for adjusting commentary content based on the identified emotion" refers to a technique for flexibly changing the generated commentary content according to the user's emotional state.

[0647] This invention is an educational system that not only solves problems that users do not understand when studying, but also recognizes the user's emotions and provides appropriate instruction methods. This system provides a series of functions: the user inputs an image of a problem, and a generative AI generates a solution and explanation, which is then output to the user. In addition, by incorporating a new emotion recognition engine, the system has the function of recognizing the user's emotions and adjusting the content of the explanation according to the user's emotions.

[0648] The main hardware and software used in implementing this invention are as follows:

[0649] Hardware

[0650] Smart Glasses (e.g., General Company Name, Smart Glasses)

[0651] server

[0652] Interface device (e.g., generic name, tablet, PC)

[0653] software

[0654] Emotion recognition engine (e.g., general company name, emotion recognition API)

[0655] Generative AI (e.g., general company name, AI generation engine)

[0656] Image recognition software using OCR technology

[0657] Voice Recognition Technology

[0658] System configuration and processing procedures

[0659] Device (smart glasses, etc.)

[0660] The user takes a picture of the problem and saves it on their device, which then sends the image data to the server.

[0661] server

[0662] The server receives the image data sent from the device, extracts text data from the image using OCR, and then formats the extracted text data into a question format and sends it to the generation AI.

[0663] Generation AI

[0664] The generation AI analyzes the received problem data, generates a solution and explanation, and sends the results back to the server.

[0665] server

[0666] The server receives the solution and explanation sent from the generated AI and sends it to the terminal, which displays the solution and explanation to the user.

[0667] User

[0668] The user checks the explanation and asks questions about parts they don't understand, and the device sends the questions to the server.

[0669] Re-explanation by generative AI

[0670] The server sends the question to the AI ​​generator, which then generates an explanation. The explanation is then sent to the device via the server and displayed to the user.

[0671] emotion recognition

[0672] The server receives the user's voice and video data and sends it to an emotion recognition engine. The emotion recognition engine analyzes the user's emotions and returns the results to the server. The server then sends the results to the generation AI, which adjusts the commentary content. The adjusted commentary is then sent to the device via the server and displayed to the user.

[0673] Specific examples

[0674] How to solve math problems

[0675] For example, a user inputs the math problem "2x + 3 = 7" into the system. The following specific prompts are fed into the generative AI model to get an explanation:

[0676] Explain in detail about "2x + 3 = 7," focusing on areas that students may find confusing. For example, explain in detail why the constant term is shifted.

[0677] In this way, the user can receive explanations in a way that is easy to understand. The system recognizes the user's emotions and adjusts the explanation content accordingly, providing effective learning support.

[0678] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0679] Step 1:

[0680] Taking and receiving images with your device

[0681] The user takes a picture of the problem they want to learn using a device such as smart glasses, and the captured image is saved on the device.

[0682] Input: An image of the problem taken by the user

[0683] Output: Image data stored in the device

[0684] Specific operation: The user uses the device's camera function to take and save the image in question.

[0685] Step 2:

[0686] Sending images from the device to the server

[0687] The terminal transmits the captured image data to the server.

[0688] Input: Image data stored on the device

[0689] Output: Image data sent to the server

[0690] Specific operation: The terminal uses its wireless communication function to upload image data to the server.

[0691] Step 3:

[0692] Image analysis by server

[0693] The server analyzes the received image and extracts the text data using OCR technology.

[0694] Input: Image data sent to the server

[0695] Output: Extracted text data

[0696] Specific operation: Using OCR technology, the text information in the image is read and formatted as problem data.

[0697] Step 4:

[0698] Generative AI generates solutions and explanations

[0699] The server sends the extracted text data to the generation AI, which then generates a solution and explanation based on this.

[0700] Input: Text data sent from the server

[0701] Output: Generated solution and explanation

[0702] Specific operation: The generative AI analyzes problem data and generates the optimal solution and accompanying explanation.

[0703] Step 5:

[0704] Sending the generated results from the server to the device

[0705] The server sends the solution and explanation received from the generating AI to the terminal.

[0706] Input: Solutions and explanations generated by the generative AI

[0707] Output: Solution and explanation sent to terminal

[0708] Specific operation: The server receives the output results of the generation AI and sends them to the terminal.

[0709] Step 6:

[0710] Displaying solutions and explanations on your device

[0711] The terminal displays the received solution and explanation to the user.

[0712] Input: Solution and explanation sent from the server

[0713] Output: The solution and explanation displayed to the user

[0714] Specific operation: The terminal displays the solution and explanation on the screen for the user to see.

[0715] Step 7:

[0716] User Questions

[0717] Users can check the explanations and ask questions about anything they don't understand by voice or text input.

[0718] Input: User question (voice or text input)

[0719] Output: User questions saved on the device

[0720] Specific operation: The user asks a question using the device's microphone or keyboard, and the device saves the question.

[0721] Step 8:

[0722] Sending questions from the device to the server

[0723] The terminal sends the user's question to the server.

[0724] Input: User questions saved on the device

[0725] Output: The query data sent to the server

[0726] Specific operation: The terminal uses its wireless communication function to send the user's question to the server.

[0727] Step 9:

[0728] Re-explanation generation using generative AI

[0729] The server sends the received question to the generation AI, which then generates an explanation again.

[0730] Input: Question data

[0731] Output: Regenerated commentary

[0732] Specific operation: The generation AI analyzes the question content and generates a re-explanation if necessary.

[0733] Step 10:

[0734] Emotion recognition by server

[0735] The server sends the user's voice and video data to an emotion recognition engine to recognize the user's emotions.

[0736] Input: User audio and video data

[0737] Output: Emotion recognition result

[0738] Specific operation: The server inputs audio and video data into the emotion recognition engine and receives the analysis results.

[0739] Step 11:

[0740] Adjustment of commentary content

[0741] The generative AI adjusts the commentary content based on the emotion recognition results.

[0742] Input: Emotion recognition results and generated commentary

[0743] Output: Adjusted commentary

[0744] Specific operation: Match the emotion recognition results and reconstruct the commentary content to correspond to the user's emotions.

[0745] Step 12:

[0746] Sending adjustment results from the server to the device

[0747] The server sends the adjusted explanation received from the generation AI to the terminal.

[0748] Input: Adjusted commentary

[0749] Output: Adjustment results sent to the device

[0750] Specific operation: The server sends the adjusted commentary content to the terminal.

[0751] Step 13:

[0752] Displaying a recap on a terminal

[0753] The terminal displays the recap to the user.

[0754] Input: Adjustment description sent from the server

[0755] Output: A restatement that is displayed to the user

[0756] Specific operation: The terminal displays the adjusted commentary content on the screen for the user to see.

[0757] Through the above processing steps, the user can receive flexible learning support that is tailored to his or her level of understanding and emotions.

[0758] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0759] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0760] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0761] [Second embodiment]

[0762] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0763] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0764] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0765] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0766] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0767] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0768] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0769] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0770] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0771] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0772] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0773] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0774] This invention is a system that solves incomprehensible problems that arise when users study and supports deeper learning. This system provides a series of functions: the user inputs an image of a problem, and a generative AI automatically generates a solution and explanation, which are then output to the user. Furthermore, the system responds to follow-up questions from the user and provides further explanations to deepen the user's understanding.

[0775] System configuration

[0776] The system configuration is as follows:

[0777] 1. Terminal

[0778] The terminal is the device through which the user captures and inputs the image in question, and can include a smartphone, tablet, or PC.

[0779] 2. Server

[0780] The server receives and processes data sent from the device, analyzes the problem using image recognition technology, and sends the data to the generating AI.

[0781] 3. Generation AI

[0782] The AI ​​generator analyzes the received problem data and generates solutions and explanations. It also has the ability to adjust the explanations according to the user's level of understanding.

[0783] Program processing explanation

[0784] The processing of the program will be explained in natural language below.

[0785] ---

[0786] Problem capture and analysis

[0787] Terminal

[0788] The device prompts the user to take a picture of the problem, which is then saved on the device.

[0789] The device sends the saved image to the server.

[0790] server

[0791] The server receives the image sent from the terminal.

[0792] The server analyzes the received image and converts it into text data using image recognition technology (OCR).

[0793] The converted text data is formatted into a question and sent to the generation AI.

[0794] ---

[0795] Generate and display solutions and explanations

[0796] Generation AI

[0797] The generation AI analyzes the problem data it receives and generates a solution and explanation.

[0798] For example, the generated solution to the problem "2x + 3 = 7" would be the steps "2x = 7 - 3," "2x = 4," and "x = 2."

[0799] server

[0800] The server receives the solution and explanation from the generated AI.

[0801] The server sends the solution and explanation to the device.

[0802] Terminal

[0803] The terminal displays the solution and explanation to the user.

[0804] ---

[0805] Handling questions and restatements

[0806] User

[0807] If the user does not understand any part of the explanation, they can ask questions using voice recognition or text input.

[0808] Terminal

[0809] The device converts the user's speech into text.

[0810] The device sends the question to the server.

[0811] server

[0812] The server analyzes the received question and sends it to the generation AI.

[0813] Generation AI

[0814] Generative AI generates new explanations based on user questions.

[0815] server

[0816] The server receives the re-explanation from the generating AI and sends it to the terminal.

[0817] Terminal

[0818] The terminal displays the explanation again to the user.

[0819] If the user has further questions, the process is repeated.

[0820] ---

[0821] Specific examples

[0822] Example of a math problem

[0823] Consider the case where a user inputs the math problem "2x + 3 = 7." The following occurs:

[0824] 1. Problem capture

[0825] The user takes a picture of the problem using the device's camera.

[0826] The device sends the captured image to the server.

[0827] 2. Image Recognition and Problem Analysis

[0828] The server analyzes the received image and uses OCR to extract the text data "2x + 3 = 7".

[0829] The extracted text data is sent to the generation AI.

[0830] 3. Generating solutions and explanations

[0831] The generative AI generates the solutions to "2x + 3 = 7": "2x = 7 - 3," "2x = 4," and "x = 2," along with their explanations.

[0832] The server sends the generated solution and explanation to the terminal.

[0833] The terminal displays the solution and explanation to the user.

[0834] 4. Questions and Restatements

[0835] The user asks, "I don't understand why we subtract 3 from 7."

[0836] The device converts the speech into text and sends it to the server.

[0837] The server sends the question to the generation AI.

[0838] The generator AI generates a restatement: "To solve the equation, we need to shift the constant term, so we subtract 3 from 7."

[0839] The server sends the re-explanation to the terminal.

[0840] The terminal displays the re-explanation to the user.

[0841] In this way, questions can be asked and re-explained as many times as necessary until the user understands. This system allows users to learn at their own pace and achieve a deep understanding.

[0842] The processing flow will be explained below.

[0843] ---

[0844] Step 1:

[0845] The user takes a picture of the problem to be studied on the terminal.

[0846] The device will display instructions to the user saying, "Please take a picture of the problem."

[0847] The user takes a picture of the problem using the device's camera.

[0848] The device saves the captured image to its internal storage.

[0849] Step 2:

[0850] The device sends the saved image to the server.

[0851] The server receives the image data sent from the terminal.

[0852] The server temporarily stores the received image data.

[0853] Step 3:

[0854] The server uses image recognition technology (OCR) to extract text data from the image.

[0855] The server formats the extracted text data into a question format.

[0856] The server sends the formatted problem data to the generation AI.

[0857] Step 4:

[0858] The generation AI analyzes the problem data it receives.

[0859] The generative AI generates a solution to the provided problem.

[0860] Generative AI generates detailed explanations based on the solution.

[0861] Step 5:

[0862] The server receives the generated solution and explanation from the generated AI.

[0863] The server sends the received solution and explanation to the terminal.

[0864] Step 6:

[0865] The terminal displays the solution and explanation to the user.

[0866] The user checks the explanation and finds parts that he does not understand.

[0867] Step 7:

[0868] The user asks the terminal questions about anything they don't understand.

[0869] In the case of voice recognition, the user speaks a question into the terminal.

[0870] For text input: The user types the question into the device.

[0871] The device converts the user's speech into text (in the case of speech recognition).

[0872] The terminal sends a question from the user to the server.

[0873] Step 8:

[0874] The server receives the user's query.

[0875] The server sends the received question to the generation AI.

[0876] Step 9:

[0877] The generative AI generates a re-explanation based on the user's question.

[0878] The generating AI sends the new explanation to the server.

[0879] Step 10:

[0880] The server receives a re-explanation from the generated AI.

[0881] The server transmits the received explanation to the terminal.

[0882] Step 11:

[0883] The terminal displays the re-explanation to the user.

[0884] The user reviews the re-explanation and determines whether or not they understand it.

[0885] If not understood, the user asks the question again (return to step 7).

[0886] ---

[0887] By repeating this series of processes, the user can continue learning until their understanding deepens.

[0888] Example 1

[0889] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0890] In today's educational environment, users often encounter points they cannot understand while studying, and a system that can appropriately address these issues is needed. In particular, users who study independently have limited access to appropriate explanations and supplementary information. Therefore, the objective of this invention is to provide a system that can immediately resolve questions users encounter during the problem-solving process and deepen their understanding.

[0891] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0892] In this invention, the server includes a means for receiving an image of a problem provided by a user, an image recognition means for analyzing the problem from the received image, and a means having a generation AI for processing the analyzed problem data, which enables a means for a user to take a picture of the problem with a camera on a terminal and send it to the server, a means for the generation AI to analyze the problem based on the prompt and automatically generate steps for a solution and explanation, and a means for converting a user's voice input into text.

[0893] A "user" is an entity that uses the system to enter problems and receive solutions and explanations.

[0894] A "problem image" is a digital image that visually represents the problem the user wants to solve.

[0895] A "terminal" is a device that a user uses to capture and input the image in question, and includes smartphones, tablets, personal computers, etc.

[0896] A "server" is the central part of the system that receives and processes data sent from the terminals.

[0897] "Image recognition means" refers to techniques and devices that extract text data from received images.

[0898] "Text data" is character information extracted by image recognition means.

[0899] "Generative AI" is artificial intelligence that generates solutions and explanations based on analyzed problem data.

[0900] A "solution" is a procedure and computational process for solving a given problem.

[0901] "Explanation" is a text that explains each step of the solution and the background and reasons behind it.

[0902] A "prompt" is a form of text or data that is given to a generative AI as instructions or input to solve a problem.

[0903] "Voice input" refers to the act of inputting voice uttered by a user as information.

[0904] "Speech recognition means" refers to the technology and devices that convert a user's voice input into text.

[0905] A "re-explanation" is a new explanation generated by the AI ​​based on additional questions from the user.

[0906] This invention is a system that solves problems that arise when users are studying and supports deepening their learning. This system provides a series of functions in which the user inputs an image of a problem, and a generative AI automatically generates a solution and explanation, which are then output to the user. In addition, the system has a mechanism that responds to follow-up questions from the user and provides further explanations to deepen the user's understanding.

[0907] System configuration

[0908] 1. Terminal

[0909] The terminal is a device that users use to take and input images of the problem, and mainly includes smartphones, tablets, and PCs. It temporarily stores the images of the problem taken by the user and then sends them to the server.

[0910] 2. Server

[0911] The server receives the image sent from the device and uses Optical Character Recognition (OCR) technology to analyze the received image. For example, it uses a library such as Tesseract OCR to extract text data from the image. The extracted text data is then sent to the generation AI.

[0912] 3. Generation AI

[0913] The generating AI analyzes the problem according to the prompts based on the text data sent from the server, and generates a solution and explanation. The generating AI can use a generative model such as GPT. The generated solution and explanation are sent to the device via the server.

[0914] 4. Voice Recognition Methods

[0915] When a user has questions about something they don't understand, they can ask using voice recognition or text input. The device converts the user's voice into text. Voice recognition can use services such as the Google Speech-to-Text API or Amazon Transcribe.

[0916] Specific examples

[0917] Consider the case where a user inputs a math problem, "2x + 3 = 7." The system operates as follows:

[0918] 1. Problem capture

[0919] The user takes the image in question using the device's camera. At this time, the device launches the smartphone's camera app and displays a "Take a Photo" button. When the user presses the "Take a Photo" button, the image is captured and temporarily saved in the device's memory. The device then sends the image data to the server.

[0920] 2. Image Recognition and Problem Analysis

[0921] The server passes the received image data to a local image analysis engine, which uses OCR technology to extract the text data "2x + 3 = 7" from the image. This text data is then sent to the generation AI.

[0922] 3. Generating solutions and explanations

[0923] The generative AI analyzes the problem based on the prompt and generates the solutions "2x = 7 - 3", "2x = 4", "x = 2" and detailed step-by-step explanations. For example, the following prompts are used:

[0924] Please solve the equation 2x + 3 = 7 and explain each step.

[0925] The generated solution and explanation are sent to the terminal via the server.

[0926] 4. Displaying the solution and explanation

[0927] The device then displays the received solutions and explanations in a user-friendly format on the screen, with each step displayed step-by-step and accompanied by a detailed explanation for that step.

[0928] 5. Questions and Restatements

[0929] The user asks, "I don't understand why 3 is subtracted from 7." The device converts the user's voice into text using speech recognition technology and sends it to the server. The server sends the question as a prompt to the generation AI, which then generates a new explanation. This new explanation is sent to the device via the server and displayed to the user again.

[0930] In this way, questions and re-explanations can be repeated as many times as necessary until the user understands. Through this system, users can learn at their own pace and achieve a deep understanding.

[0931] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0932] Step 1: Take a picture of the problem and send it to us

[0933] Terminal

[0934] The user takes the image in question using the device's camera. For example, they launch the camera app on their smartphone and a "Take Photo" button appears.

[0935] When the user presses the "take a photo" button, an image is captured and temporarily saved in the device's internal memory.

[0936] The device sends the stored image data to the server using an HTTP POST request.

[0937] Input: A user-taken image of the problem

[0938] Output: Image data sent to the server

[0939] Specific behavior:

[0940] 1. The user activates the device camera.

[0941] 2. Press the "Capture" button to capture the image in question.

[0942] 3. Temporarily save image data.

[0943] 4. Send the image data to the server.

[0944] Step 2: Receiving and analyzing images

[0945] server

[0946] The server saves the image data received via the HTTP request in local storage.

[0947] The server passes the saved image data to an OCR (Optical Character Recognition) engine and begins processing.

[0948] OCR technology (e.g., Tesseract OCR) is used to extract text data from the image.

[0949] The extracted text data is formatted into questions and prompts are created to be sent to the generation AI.

[0950] Input: Image data sent from the device

[0951] Output: Problem text data to send to the generation AI

[0952] Specific behavior:

[0953] 1. The server receives the image data.

[0954] 2. Save the received image data locally.

[0955] 3. Pass the image data to the OCR engine.

[0956] 4. Extract text data using OCR technology.

[0957] 5. Format the text data into a question format.

[0958] 6. Create a prompt to send the formatted text data to the generation AI.

[0959] Step 3: Generate solutions and explanations

[0960] Generation AI

[0961] The generation AI receives the text data sent from the server.

[0962] The system analyzes the received text data as prompts and automatically generates steps to solve the problem and detailed explanations for each.

[0963] The generated solutions are in the form of, for example, "2x = 7 - 3," "2x = 4," or "x = 2." The explanation is something like, "First, subtract 3 on the left side from 7 on the right side."

[0964] Input: Question text data sent from the server

[0965] Output: Generated solution and explanation

[0966] Specific behavior:

[0967] 1. The generation AI receives the text data.

[0968] 2. Analyze text data as prompts.

[0969] 3. Automatically generate solutions and detailed explanations.

[0970] Step 4: Submit and view your solution and explanation

[0971] server

[0972] The server receives the generated solution and explanation from the generating AI.

[0973] The integrity of the received data is checked and it is sent to the terminal in JSON format or similar.

[0974] Input: Solution and explanation sent from the generating AI

[0975] Output: Solution and explanation data sent to the terminal

[0976] Specific behavior:

[0977] 1. The server receives the solution and explanation.

[0978] 2. Check data integrity.

[0979] 3. The data whose integrity has been confirmed is sent to the terminal.

[0980] Terminal

[0981] The terminal analyzes the solutions and explanations received from the server and displays them on the screen in a format that is easy for the user to see.

[0982] Each step is displayed step by step and a detailed explanation for each step is also provided.

[0983] Input: Solution and explanation data sent from the server

[0984] Output: The solution and explanation displayed to the user

[0985] Specific behavior:

[0986] 1. The device receives the solution and explanation.

[0987] 2. The solution and explanation are analyzed and displayed on the screen.

[0988] Step 5: Submit your question and voice input

[0989] User

[0990] If the user does not understand the explanation, they can enter a question using voice recognition or text input.

[0991] For example, tap the microphone icon on your smartphone to start voice input and say, "I don't know why I subtract 3 from 7."

[0992] Input: Additional questions about the explanation

[0993] Output: Input question data to the terminal

[0994] Specific behavior:

[0995] 1. The user taps the microphone icon.

[0996] 2. The user types in a question by voice.

[0997] Terminal

[0998] The device converts the user's speech into text using services such as the Google Speech-to-Text API or Amazon Transcribe.

[0999] The converted text data is sent to the server.

[1000] Input: A spoken question from the user

[1001] Output: Sends text-formatted question data to the server

[1002] Specific behavior:

[1003] 1. The device receives the audio data.

[1004] 2. Convert voice to text using voice recognition technology.

[1005] 3. Send the text data to the server.

[1006] Step 6: Parse the question and generate a restatement

[1007] server

[1008] The server receives the question sent from the terminal.

[1009] The received text data is sent to the generation AI as a prompt.

[1010] Input: Text question data sent from the terminal

[1011] Output: Prompt data to send to the generation AI

[1012] Specific behavior:

[1013] 1. The server receives the question text data.

[1014] 2. Send the text data to the generation AI as a prompt.

[1015] Generation AI

[1016] Generative AI generates new explanations based on user questions.

[1017] For example, it generates a restatement like "To solve the equation, we need to shift the constant term, so subtract 3 from 7."

[1018] Input: Question prompt sent from the server

[1019] Output: Generated restatement

[1020] Specific behavior:

[1021] 1. The generative AI receives a question prompt.

[1022] 2. Automatically generate new explanations based on questions.

[1023] Step 7: Submit and view the recap

[1024] server

[1025] The server receives the generated re-explanation from the generation AI and sends it to the terminal.

[1026] Input: Restatement sent by the generating AI

[1027] Output: Recap what is sent to the terminal

[1028] Specific behavior:

[1029] 1. The server receives the re-explanation.

[1030] 2. Send the re-explanation to the device.

[1031] Terminal

[1032] The terminal displays the re-explanation received from the server to the user.

[1033] Input: Recap sent by server

[1034] Output: A restatement that is displayed to the user

[1035] Specific behavior:

[1036] 1. The device receives the re-explanation.

[1037] 2. Display the re-explanation to the user.

[1038] In this way, the user can ask questions and receive re-explanations as many times as necessary until they understand. This system allows users to learn at their own pace and achieve a deep understanding.

[1039] (Application example 1)

[1040] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1041] In factories, work procedures are complex, and employees often make mistakes, leading to a decline in production efficiency and the inability to quickly resolve problems that arise during the work process. Newly introduced manuals and work procedures also tend to be poorly understood. A system is needed that allows employees to easily understand work procedures and receive assistance in resolving problems.

[1042] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1043] In this invention, the server includes means for receiving an image of a task provided by a user, image recognition means for analyzing the task from the received image, means having a generation AI for processing the analyzed task data, means for outputting to the user a solution and explanation obtained by the generation AI, means for receiving questions from the user, means for generating an explanation based on the received question and outputting it to the user, and means for displaying the generated solution and procedure through a smart device to support complex work procedures and problem solving in a factory. This enables employees to quickly solve problems they encounter during work and work efficiently.

[1044] The "means for receiving images of tasks provided by users" includes a device that allows a factory employee to take an image of a work procedure manual, a drawing, etc., and transmit the image to a server.

[1045] "Image recognition means" refers to technology that analyzes received images, identifies and extracts text and figures from the images, and converts them into data.

[1046] "Means having generative AI" refers to an artificial intelligence (AI) system that analyzes input data and automatically generates solutions and explanations.

[1047] "Means for outputting solutions and explanations to users" refers to technology for providing the solutions and explanations generated by the generation AI to factory employees via display devices or audio output devices.

[1048] The term "means for receiving questions" refers to technology that provides an interface for factory employees to input or voice questions and receives those questions.

[1049] "Means of generating new explanations based on questions and outputting them to users" refers to technology that uses a generative AI to reanalyze received questions, generate new explanations and procedures, and provide them to factory employees.

[1050] "Means for displaying generated solutions and procedures through smart devices to support complex work procedures and problem-solving within factories" refers to technology that uses wearable devices such as smart glasses or tablet terminals to display generated procedures and solutions to factory employees in real time.

[1051] The present invention is a system for supporting complex work procedures and problem solving within a factory, and in particular provides work procedures and solutions to factory employees using wearable devices such as smart glasses and tablet terminals.

[1052] System configuration

[1053] The system configuration is as follows:

[1054] 1. Terminal

[1055] Terminals are devices that factory employees use to take and input images of work procedures and drawings. These devices include smart glasses and tablets. These terminals are equipped with a camera function and send the images to a server.

[1056] 2. Server

[1057] The server is responsible for receiving and processing data sent from the terminal. The server has the following functions:

[1058] Image recognition means: The received image is analyzed and converted into text data using OCR (optical character recognition) technology. This converted data is sent to the generation AI.

[1059] Question analysis means: Receives and analyzes questions from factory employees. This data is also sent to the generation AI.

[1060] 3. Generation AI

[1061] The generative AI analyzes the received problem data and questions, and generates solutions and explanations. It also has the ability to adjust the explanation content according to the user's level of understanding. Examples of generative AI models used include GPT-4.

[1062] Program processing explanation

[1063] Problem capture and analysis

[1064] 1. The device instructs factory employees to take pictures of work procedures and drawings, which are then saved on the device.

[1065] 2. The device sends the saved image to the server.

[1066] 3. The server receives the image sent from the device.

[1067] 4. The server analyzes the received image and converts it into text data using OCR.

[1068] 5. The converted text data is formatted into a question and sent to the generation AI.

[1069] Generate and display solutions and explanations

[1070] 1. The generation AI analyzes the problem data it receives and generates specific work procedures and solutions.

[1071] 2. The generated procedures and solutions are sent to the terminal via the server.

[1072] 3. The terminal displays the generated solution and procedures to the factory worker.

[1073] Handling questions and restatements

[1074] 1. If a factory employee has questions about parts of the explanation that they don't understand, they can ask using voice recognition or text input.

[1075] 2. The device converts the user's speech into text.

[1076] 3. The device sends the question to the server.

[1077] 4. The server analyzes the received question and sends it to the generation AI.

[1078] 5. Generative AI generates new explanations based on the user's questions.

[1079] 6. The server receives the re-explanation from the generation AI and sends it to the device.

[1080] 7. The device displays the explanation again to the user.

[1081] Specific examples

[1082] Work procedure support example

[1083] Consider a case where a factory worker does not understand the work procedure "attach part A to part B." The following process is performed:

[1084] 1. A factory employee uses the device's camera to take an image of a work procedure manual or drawing.

[1085] 2. The device sends the captured image to the server.

[1086] 3. The server analyzes the received image and extracts text data using OCR.

[1087] 4. Send the extracted text data to the generation AI.

[1088] 5. The generative AI generates a solution and explanation such as "Steps for assembling parts A and B: 1. Remove the screws from part A. 2. Attach part A to part B."

[1089] 6. The generated procedure is sent to the terminal via the server and displayed to the factory employee.

[1090] 7. A factory worker asks, "Which way do you turn the screw?"

[1091] 8. The question is converted into text using voice recognition technology and sent to the server.

[1092] 9. The server sends the question to the generation AI, which generates a re-explanation: "Rotate clockwise."

[1093] 10. The re-explanation is sent to the terminal and displayed to the factory employee.

[1094] As a result, the system of the present invention can improve the work efficiency of users in a factory and speed up the problem solving process.

[1095] Prompt Sentence Examples

[1096] "Please solve the following problem: I don't understand step C: "Step C: Attach part A to part B.""

[1097] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1098] Step 1:

[1099] The terminal instructs the factory employee to take an image of the work procedure manual or blueprint. The factory employee takes an image of the work procedure manual or blueprint using the camera on the terminal. The input of this step is the factory employee's photographing behavior and the captured image, and the output is an image file saved on the terminal.

[1100] Step 2:

[1101] The device sends the saved image to the server. The device sends the image file taken by the factory employee to the server using an HTTP POST request. The input of this step is the saved image file, and the output is the image data received by the server.

[1102] Step 3:

[1103] The server receives the image sent from the terminal. The server temporarily stores the received image data. The input of this step is the image data sent from the terminal, and the output is the image data stored on the server.

[1104] Step 4:

[1105] The server analyzes the received image and converts it into text data using OCR. The server uses OCR technology (e.g., pytesseract) to extract text information from the image. The input for this step is the image data stored on the server, and the output is the extracted text data.

[1106] Step 5:

[1107] The server formats the converted text data into a question format and sends it to the generation AI. The server formats the text data into an appropriate format (e.g., JSON format) and sends an API request to the generation AI. The input of this step is the extracted text data, and the output is the question format data sent to the generation AI.

[1108] Step 6:

[1109] The generative AI analyzes the received problem data and generates a specific work procedure or solution. The generative AI (e.g., GPT-4) generates a solution and explanation based on the problem data. The input of this step is data in the form of a problem, and the output is the generated solution and explanation.

[1110] Step 7:

[1111] The server sends the generated solution and explanation to the terminal. The server sends the data received from the generation AI to the terminal. The input of this step is the generated solution and explanation, and the output is the solution and explanation sent to the terminal.

[1112] Step 8:

[1113] The terminal displays the generated solution and explanation to the factory employee. The terminal displays the received solution and explanation on the display. The input of this step is the solution and explanation sent to the terminal, and the output is display information that can be viewed by the factory employee.

[1114] Step 9:

[1115] If a factory employee does not understand a part of the explanation, they ask a question using voice recognition or text input. The factory employee then uses the terminal's voice input function or keyboard to input the question. The input for this step is the factory employee's voice or text input, and the output is the question data recognized by the terminal.

[1116] Step 10:

[1117] The terminal converts the user's voice into text. The terminal converts the voice data into text using voice recognition technology. The input of this step is the voice data of the factory employee, and the output is text data.

[1118] Step 11:

[1119] The terminal sends the question content to the server. The terminal then sends the question content converted into text data to the server. The input to this step is the text data, and the output is the question data sent to the server.

[1120] Step 12:

[1121] The server analyzes the received question and sends it to the generation AI. The server then formats the question and sends an API request to the generation AI again. The input of this step is the question data received by the server, and the output is the question data sent to the generation AI.

[1122] Step 13:

[1123] The generation AI generates a new explanation based on the user's question. The generation AI analyzes the question data and generates an appropriate re-explanation. The input of this step is the question data sent to the generation AI, and the output is the generated re-explanation data.

[1124] Step 14:

[1125] The server receives the re-explanation from the generation AI and sends it to the terminal. The server sends the generated re-explanation data to the terminal. The input of this step is the generated re-explanation data, and the output is the re-explanation data sent to the terminal.

[1126] Step 15:

[1127] The terminal displays the re-explanation to the factory employee. The terminal displays the re-explanation data on a display and provides it to the factory employee. The input of this step is the re-explanation data sent to the terminal, and the output is re-explanation information that can be viewed by the factory employee.

[1128] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1129] This invention is an educational system that not only solves problems that users do not understand when studying, but also recognizes the user's emotions and provides appropriate instruction methods. This system provides a series of functions: the user inputs an image of a problem, and a generative AI generates a solution and explanation, which is then output to the user. Furthermore, by incorporating a new emotion engine, the system has the function of recognizing the user's emotions and adjusting the explanation content according to the user's emotions.

[1130] System configuration

[1131] The system configuration is as follows:

[1132] 1. Terminal

[1133] The terminal is the device through which the user captures and inputs the image in question, and can include smartphones, tablets, and personal computers.

[1134] 2. Server

[1135] The server receives and processes data sent from the device, analyzes the problem using image recognition technology, and sends the data to the generating AI.

[1136] 3. Generation AI

[1137] The AI ​​generator analyzes the received problem data and generates solutions and explanations. It also has the ability to adjust the explanations according to the user's level of understanding.

[1138] 4. Emotion Engine

[1139] The emotion engine recognizes the user's emotions and provides a means for the system to respond based on those emotions. Specifically, it analyzes voice and facial expressions to determine whether the user understands or is confused.

[1140] Program processing explanation

[1141] The processing of the program will be explained in natural language below.

[1142] ---

[1143] Problem capture and analysis

[1144] Terminal

[1145] The device prompts the user to take a picture of the problem, which is then saved on the device.

[1146] The user takes a picture of the problem using the device's camera function.

[1147] The device sends the saved image to the server.

[1148] server

[1149] The server receives the image data sent from the terminal.

[1150] The server temporarily stores the received image data.

[1151] The server uses image recognition technology (OCR) to extract text data from the image.

[1152] The extracted text data is formatted into a question and sent to the generation AI.

[1153] ---

[1154] Generate and display solutions and explanations

[1155] Generation AI

[1156] The generation AI analyzes the problem data it receives.

[1157] The generative AI generates a solution to the provided problem.

[1158] Generative AI generates detailed explanations based on the solution.

[1159] server

[1160] The server receives the generated solution and explanation from the generated AI.

[1161] The server sends the received solution and explanation to the terminal.

[1162] Terminal

[1163] The terminal displays the solution and explanation to the user.

[1164] The user checks the explanation and finds parts that he does not understand.

[1165] ---

[1166] Handling questions and restatements

[1167] User

[1168] The user asks the terminal questions about anything they don't understand.

[1169] In the case of voice recognition, the user speaks a question into the terminal.

[1170] For text input: The user types the question into the device.

[1171] Terminal

[1172] The device converts the user's speech into text (in the case of speech recognition).

[1173] The terminal sends a question from the user to the server.

[1174] server

[1175] The server receives the user's query.

[1176] The server sends the received question to the generation AI.

[1177] Generation AI

[1178] The generative AI generates a re-explanation based on the user's question.

[1179] The generating AI sends the new explanation to the server.

[1180] server

[1181] The server receives a re-explanation from the generated AI.

[1182] The server transmits the received explanation to the terminal.

[1183] Terminal

[1184] The terminal displays the re-explanation to the user.

[1185] The user reviews the re-explanation and determines whether or not they understand it.

[1186] If the user does not understand, he or she asks the question again.

[1187] ---

[1188] Recognizing and Responding to Emotions

[1189] Terminal

[1190] The device collects the user's voice and camera footage.

[1191] The terminal transmits the collected data to the server.

[1192] server

[1193] The server sends the received data to the emotion engine.

[1194] The emotion engine analyzes voice and facial expressions to recognize the user's emotions.

[1195] The emotion engine sends the recognition results to the server.

[1196] Generation AI

[1197] The generative AI adjusts the commentary content based on the recognition results.

[1198] For example, if the user is confused, generate a more detailed explanation.

[1199] server

[1200] The server receives the adjusted commentary from the generated AI.

[1201] The server sends the adjusted commentary to the terminal.

[1202] Terminal

[1203] The terminal displays the adjusted description to the user.

[1204] ---

[1205] Specific examples

[1206] Example of a math problem

[1207] Consider the case where a user inputs the math problem "2x + 3 = 7." The following occurs:

[1208] 1. Problem capture

[1209] The user takes a picture of the problem using the device's camera.

[1210] The device sends the captured image to the server.

[1211] 2. Image Recognition and Problem Analysis

[1212] The server analyzes the received image and uses OCR to extract the text data "2x + 3 = 7".

[1213] The extracted text data is sent to the generation AI.

[1214] 3. Generating solutions and explanations

[1215] The generative AI generates the solutions to "2x + 3 = 7": "2x = 7 - 3," "2x = 4," and "x = 2," along with their explanations.

[1216] The server sends the generated solution and explanation to the terminal.

[1217] The terminal displays the solution and explanation to the user.

[1218] 4. Questions and Restatements

[1219] The user asks, "I don't understand why we subtract 3 from 7."

[1220] The device converts the speech into text and sends it to the server.

[1221] The server sends the question to the generation AI.

[1222] The generator AI generates a restatement: "To solve the equation, we need to shift the constant term, so we subtract 3 from 7."

[1223] The server sends the re-explanation to the terminal.

[1224] The terminal displays the re-explanation to the user.

[1225] 5. Emotion Recognition and Response

[1226] The device collects the user's voice and camera footage and sends them to the server.

[1227] The emotion engine analyzes voice and facial expressions to recognize the user's emotions.

[1228] The generative AI adjusts the commentary based on the perceived emotion (generating a more detailed explanation if the person is confused).

[1229] The server sends the adjusted commentary to the terminal, which displays it to the user.

[1230] In this way, users can receive support tailored to their level of understanding and emotional state. This system allows users to learn at their own pace and achieve deep understanding.

[1231] The processing flow will be explained below.

[1232] I understand. Now, I will explain the process flow of the invention that combines the emotion engine in concrete steps.

[1233] ---

[1234] Step 1:

[1235] The device will prompt the user to "Take a picture of the problem."

[1236] The user uses the device's camera to take the image in question.

[1237] The device saves the captured image to its internal storage.

[1238] Step 2:

[1239] The device sends the saved image to the server.

[1240] The server receives the image data sent from the terminal.

[1241] The server temporarily stores the received image data.

[1242] Step 3:

[1243] The server uses image recognition technology (OCR) to extract text data from the image.

[1244] The server formats the extracted text data into a question format.

[1245] The server sends the formatted problem data to the generation AI.

[1246] Step 4:

[1247] The generation AI analyzes the problem data it receives.

[1248] The generative AI generates a solution to the provided problem.

[1249] Generative AI generates detailed explanations based on the solution.

[1250] Step 5:

[1251] The server receives the generated solution and explanation from the generated AI.

[1252] The server sends the received solution and explanation to the terminal.

[1253] Step 6:

[1254] The terminal displays the solution and explanation to the user.

[1255] The user checks the explanation and finds parts that he does not understand.

[1256] Step 7:

[1257] The user asks the terminal questions about anything they don't understand.

[1258] In the case of voice recognition, the user speaks a question into the terminal.

[1259] For text input: The user types the question into the device.

[1260] Step 8:

[1261] The device converts the user's speech into text (in the case of speech recognition).

[1262] The terminal sends a question from the user to the server.

[1263] Step 9:

[1264] The server receives the user's query.

[1265] The server sends the received question to the generation AI.

[1266] Step 10:

[1267] The generative AI generates a re-explanation based on the user's question.

[1268] The generating AI sends the new explanation to the server.

[1269] Step 11:

[1270] The server receives a re-explanation from the generated AI.

[1271] The server transmits the received explanation to the terminal.

[1272] Step 12:

[1273] The terminal displays the re-explanation to the user.

[1274] The user reviews the re-explanation and determines whether or not they understand it.

[1275] If not understood, the user asks the question again (return to step 7).

[1276] Step 13:

[1277] The device collects the user's voice and camera footage.

[1278] The terminal transmits the collected data to the server.

[1279] Step 14:

[1280] The server sends the received data to the emotion engine.

[1281] The emotion engine analyzes voice and facial expressions to recognize the user's emotions.

[1282] The emotion engine sends the recognition results to the server.

[1283] Step 15:

[1284] The generative AI adjusts the commentary content based on the recognition results.

[1285] For example, if the user is confused, generate a more detailed explanation.

[1286] Step 16:

[1287] The server receives the adjusted commentary from the generated AI.

[1288] The server sends the adjusted commentary to the terminal.

[1289] Step 17:

[1290] The terminal displays the adjusted description to the user.

[1291] ---

[1292] Through this series of processes, users can receive support tailored to their level of understanding and emotional state.

[1293] Example 2

[1294] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1295] In recent years, there has been a growing need for educational support systems that can quickly and accurately resolve user issues. However, conventional systems lack the explanations and suggestions needed to resolve issues that users cannot understand. Furthermore, because they do not take the user's emotional state into consideration, learning effectiveness may not be fully realized. To solve these problems, there is a need for the development of a system that not only provides answers to users' questions but also recognizes the user's emotions and provides appropriate guidance accordingly.

[1296] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1297] In this invention, the server includes means for receiving images of problems provided by the user, image recognition means for analyzing the problems from the received images, artificial intelligence means for processing the analyzed problem data, means for outputting to the user the solutions and explanations obtained by the artificial intelligence means, means for receiving questions from the user, means for generating new explanations based on the received questions and outputting them to the user, emotion recognition means for collecting and analyzing emotion data of the user, and means for adjusting the content of the explanations based on the analysis results. This makes it possible to solve the user's understanding problems and provide appropriate guidance according to the user's emotional state, thereby maximizing the learning effect.

[1298] "User" refers to a person who uses the system.

[1299] "Question image" refers to an image file containing a study question that the user has taken using the camera function.

[1300] "Means for receiving" refers to the function that allows the system to receive problem images and questions provided by the user.

[1301] "Image recognition means" refers to technology (e.g., OCR technology) for extracting text data or problem information from received images.

[1302] "Generative AI" refers to a program that generates solutions and explanations from problem data analyzed using artificial intelligence technology.

[1303] "Means of output" refers to functions for providing users with the solutions and explanations obtained by the generating AI, such as screen display and audio output functions.

[1304] "Means for receiving questions" refers to a function for receiving additional questions or inquiries from users.

[1305] "Means for generating a new explanation" refers to the function of the generation AI to create a new explanation in response to a question from the user.

[1306] "Emotion data" refers to information about emotions analyzed from the user's voice, video, etc.

[1307] "Emotion recognition means" refers to technology for analyzing collected emotion data and recognizing the user's emotional state.

[1308] The "adjustment means" refers to a function for appropriately changing the commentary content based on the analysis results of the emotion recognition means.

[1309] This invention is an educational support system that not only resolves points that a user does not understand when studying, but also recognizes the user's emotions and provides appropriate teaching methods. This system is a combination of various hardware and software and consists of the following components:

[1310] First, the devices used by users include devices such as smartphones, tablets, and PCs. Users can use these devices to take pictures of study questions. The devices use their camera functions to acquire the images of the questions taken by the users.

[1311] The acquired image data is then sent to a server via a network. The server stores the received image data and uses image recognition technology (OCR = Optical Character Recognition) to extract text data from the image. This text data is then analyzed to determine the content of the problem.

[1312] The extracted text data is sent to the generative AI, which analyzes the problem based on the text data and generates a solution and a detailed explanation. The generative AI uses the latest machine learning and natural language processing technologies, enabling advanced analysis and generation.

[1313] The generated solutions and explanations are then sent to the terminal via the server. The terminal displays the received solutions and explanations to the user. The user can check the solutions and explanations displayed and ask additional questions about any parts they do not understand.

[1314] A user's question is entered into the device through text input or voice input. The device that receives the question sends the data to the server. The server then sends the user's question to the generation AI, which then generates a re-explanation. The generated re-explanation is sent to the device via the server and displayed to the user.

[1315] The system also incorporates emotion recognition functionality. The device collects the user's voice and camera footage and sends it to a server. The server then uses an emotion engine to analyze the user's emotion data and identify their emotional state. The emotion recognition data is then sent to a generation AI, which then adjusts the commentary content according to the user's emotional state. For example, if the user is confused, a more detailed explanation will be generated.

[1316] As a concrete example, consider the case where a user inputs the math problem "2x + 3 = 7." The user takes an image of the problem using the device's camera and sends it from the device to the server. The server uses OCR technology to extract the text data "2x + 3 = 7" and sends it to the generation AI. The generation AI generates a solution and detailed explanation for "2x + 3 = 7" and sends it to the device via the server. The device then displays the solution and explanation to the user.

[1317] If a user asks, "I don't know why we subtract 3 from 7," the device sends the question to the server, and the generation AI generates a restatement, saying, "Because we need to shift the constant term to solve the equation." The server sends the restatement to the device, which displays it to the user.

[1318] Here are some examples of specific prompts:

[1319] Problem: Generate a solution and explanation for "2x + 3 = 7".

[1320] User Question: "I don't understand why we subtract 3 from 7."

[1321] User sentiment: Confused

[1322] Please provide a more detailed explanation.

[1323] This system allows users to receive learning support at their own pace, and by receiving guidance tailored to their emotional state, they can achieve a deeper understanding.

[1324] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1325] Step 1:

[1326] The device instructs the user to take a picture of the problem. Specifically, the device displays "Please take a picture of the problem." Input: User's instruction appears. Output: Message is displayed.

[1327] Step 2:

[1328] The user takes a picture of the problem using the device's camera. Input: The image of the problem taken by the user. Output: Image data.

[1329] Step 3:

[1330] The image captured by the device is sent to the server. Input: Image data from the device. Output: Image file sent to the server.

[1331] Step 4:

[1332] The server receives the image data. Input: Image data sent from the terminal. Output: Reception confirmation signal and storage of the image data.

[1333] Step 5:

[1334] The server uses image recognition technology (OCR) to extract text data. The server launches image analysis software to extract text data from the image. Input: Received image data. Output: Extracted text data.

[1335] Step 6:

[1336] The server sends the extracted text data to the generation AI. The server calls the generation AI's API and sends the text data. Input: Text data. Output: Data sent to the generation AI.

[1337] Step 7:

[1338] The generative AI generates a solution and explanation for the problem. The generative AI analyzes the input text data and generates an appropriate solution and detailed explanation. Input: Submitted text data. Output: Generated solution and explanation data.

[1339] Step 8:

[1340] The server sends the generated solution and explanation to the terminal. The server transfers the solution and explanation data received from the generating AI to the terminal. Input: Solution and explanation data from the generating AI. Output: Sending the solution and explanation data to the terminal.

[1341] Step 9:

[1342] The terminal displays the solution and explanation to the user. The solution and explanation are displayed on the terminal screen. Input: Solution and explanation data. Output: Display to user.

[1343] Step 10:

[1344] The user inputs a question. The user checks the solution and explanation and inputs a question about the part they don't understand. Voice input or text input is available. Input: User's question (text or voice). Output: Question data.

[1345] Step 11:

[1346] The terminal sends a question to the server. Input: Question data. Output: Question data sent to the server.

[1347] Step 12:

[1348] The server sends the question to the generation AI. The server calls the generation AI's API and sends the question data. Input: User's question data. Output: Sending question data to the generation AI.

[1349] Step 13:

[1350] The generation AI generates a restatement based on the question. The generation AI analyzes the question data and generates a restatement for the question. Input: Submitted question data. Output: Generated restatement data.

[1351] Step 14:

[1352] The server sends the re-explanation to the terminal. The server transfers the re-explanation data received from the generating AI to the terminal. Input: Re-explanation data from the generating AI. Output: Sending the re-explanation data to the terminal.

[1353] Step 15:

[1354] The terminal displays the re-explanation to the user. The re-explanation is displayed on the terminal screen. Input: Re-explanation data. Output: Re-display to user.

[1355] Step 16:

[1356] The device collects the user's emotional data and sends it to the server. The device collects the user's audio and video data and sends it to the server. Input: User's audio and video data. Output: Sending emotional data to the server.

[1357] Step 17:

[1358] The server sends emotion data to the emotion engine. Input: Emotion data sent from the device. Output: Data sent to the emotion engine.

[1359] Step 18:

[1360] The emotion engine analyzes the user's emotions. The emotion engine analyzes voice and facial expressions to recognize the user's emotional state. Input: Emotion data. Output: Analyzed emotional state data.

[1361] Step 19:

[1362] The emotion engine sends the analysis results to the generation AI. The emotion engine sends the analysis results to the generation AI, which adjusts the commentary content. Input: Analysis result data. Output: Data sent to the generation AI.

[1363] Step 20:

[1364] The generation AI adjusts the commentary content according to the emotion. The generation AI adjusts the commentary content based on the recognition results and generates an appropriate commentary. Input: Analysis result data. Output: Adjusted commentary data.

[1365] Step 21:

[1366] The server sends the adjusted commentary to the terminal. The server transfers the adjusted commentary data received from the generation AI to the terminal. Input: Adjusted commentary data from the generation AI. Output: Sending adjusted commentary data to the terminal.

[1367] Step 22:

[1368] The terminal displays the adjusted commentary to the user. The adjusted commentary is displayed on the terminal screen. Input: Adjusted commentary data. Output: Adjusted commentary displayed to the user.

[1369] (Application example 2)

[1370] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1371] When a user encounters a point in the educational process that they do not understand, a common educational system has difficulty quickly recognizing that confusion and responding effectively. In particular, flexible responses based on the learner's emotions and level of understanding are required, but this has been difficult to achieve with conventional systems. In addition, there is a lack of means to grasp the learner's progress in real time and provide appropriate feedback on the spot.

[1372] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1373] In this invention, the server includes means for receiving an image of a problem provided by a user, image recognition means for analyzing the problem from the received image, means having a generation AI for processing the analyzed problem data, means for outputting to the user a solution and explanation obtained by the generation AI, means for receiving a question from the user, means for generating an explanation based on the received question and outputting it to the user, means for identifying the user's emotion, and means for adjusting the content of the explanation based on the identified emotion. This enables flexible learning support according to the user's emotion and level of understanding, and also enables real-time progress confirmation and feedback.

[1374] "User" refers to the subject who uses this system to study, and who performs operations such as entering questions and asking questions.

[1375] "Means for receiving problem images" refers to an interface for receiving problem images photographed or scanned by a user and importing them into the system.

[1376] "Image recognition means" refers to technology that analyzes the received image of a question and extracts the question content, such as text data.

[1377] "Generative AI" refers to artificial intelligence technology that generates solutions and explanations based on extracted problem data.

[1378] "Means for outputting to the user" refers to an interface for visually or audibly conveying the solution and explanation obtained by the generative AI to the user.

[1379] The "means for receiving a question" refers to an interface for receiving a question within the system when a user inputs a question about a point that he or she does not understand.

[1380] The "means for generating a new explanation and outputting it to the user" refers to an interface for generating a new explanation based on a received user question and conveying that explanation to the user visually or audibly.

[1381] "Means for identifying user emotions" refers to technology that analyzes the user's facial expressions and voice to identify the emotion the user is currently feeling.

[1382] The "means for adjusting commentary content based on the identified emotion" refers to a technique for flexibly changing the generated commentary content according to the user's emotional state.

[1383] This invention is an educational system that not only solves problems that users do not understand when studying, but also recognizes the user's emotions and provides appropriate instruction methods. This system provides a series of functions: the user inputs an image of a problem, and a generative AI generates a solution and explanation, which is then output to the user. In addition, by incorporating a new emotion recognition engine, the system has the function of recognizing the user's emotions and adjusting the content of the explanation according to the user's emotions.

[1384] The main hardware and software used in implementing this invention are as follows:

[1385] Hardware

[1386] Smart Glasses (e.g., General Company Name, Smart Glasses)

[1387] server

[1388] Interface device (e.g., generic name, tablet, PC)

[1389] software

[1390] Emotion recognition engine (e.g., general company name, emotion recognition API)

[1391] Generative AI (e.g., general company name, AI generation engine)

[1392] Image recognition software using OCR technology

[1393] Voice Recognition Technology

[1394] System configuration and processing procedures

[1395] Device (smart glasses, etc.)

[1396] The user takes a picture of the problem and saves it on their device, which then sends the image data to the server.

[1397] server

[1398] The server receives the image data sent from the device, extracts text data from the image using OCR, and then formats the extracted text data into a question format and sends it to the generation AI.

[1399] Generation AI

[1400] The generation AI analyzes the received problem data, generates a solution and explanation, and sends the results back to the server.

[1401] server

[1402] The server receives the solution and explanation sent from the generated AI and sends it to the terminal, which displays the solution and explanation to the user.

[1403] User

[1404] The user checks the explanation and asks questions about parts they don't understand, and the device sends the questions to the server.

[1405] Re-explanation by generative AI

[1406] The server sends the question to the AI ​​generator, which then generates an explanation. The explanation is then sent to the device via the server and displayed to the user.

[1407] emotion recognition

[1408] The server receives the user's voice and video data and sends it to an emotion recognition engine. The emotion recognition engine analyzes the user's emotions and returns the results to the server. The server then sends the results to the generation AI, which adjusts the commentary content. The adjusted commentary is then sent to the device via the server and displayed to the user.

[1409] Specific examples

[1410] How to solve math problems

[1411] For example, a user inputs the math problem "2x + 3 = 7" into the system. The following specific prompts are fed into the generative AI model to get an explanation:

[1412] Explain in detail about "2x + 3 = 7," focusing on areas that students may find confusing. For example, explain in detail why the constant term is shifted.

[1413] In this way, the user can receive explanations in a way that is easy to understand. The system recognizes the user's emotions and adjusts the explanation content accordingly, providing effective learning support.

[1414] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1415] Step 1:

[1416] Taking and receiving images with your device

[1417] The user takes a picture of the problem they want to learn using a device such as smart glasses, and the captured image is saved on the device.

[1418] Input: An image of the problem taken by the user

[1419] Output: Image data stored in the device

[1420] Specific operation: The user uses the device's camera function to take and save the image in question.

[1421] Step 2:

[1422] Sending images from the device to the server

[1423] The terminal transmits the captured image data to the server.

[1424] Input: Image data stored on the device

[1425] Output: Image data sent to the server

[1426] Specific operation: The terminal uses its wireless communication function to upload image data to the server.

[1427] Step 3:

[1428] Image analysis by server

[1429] The server analyzes the received image and extracts the text data using OCR technology.

[1430] Input: Image data sent to the server

[1431] Output: Extracted text data

[1432] Specific operation: Using OCR technology, the text information in the image is read and formatted as problem data.

[1433] Step 4:

[1434] Generative AI generates solutions and explanations

[1435] The server sends the extracted text data to the generation AI, which then generates a solution and explanation based on this.

[1436] Input: Text data sent from the server

[1437] Output: Generated solution and explanation

[1438] Specific operation: The generative AI analyzes problem data and generates the optimal solution and accompanying explanation.

[1439] Step 5:

[1440] Sending the generated results from the server to the device

[1441] The server sends the solution and explanation received from the generating AI to the terminal.

[1442] Input: Solutions and explanations generated by the generative AI

[1443] Output: Solution and explanation sent to terminal

[1444] Specific operation: The server receives the output results of the generation AI and sends them to the terminal.

[1445] Step 6:

[1446] Displaying solutions and explanations on your device

[1447] The terminal displays the received solution and explanation to the user.

[1448] Input: Solution and explanation sent from the server

[1449] Output: The solution and explanation displayed to the user

[1450] Specific operation: The terminal displays the solution and explanation on the screen for the user to see.

[1451] Step 7:

[1452] User Questions

[1453] Users can check the explanations and ask questions about anything they don't understand by voice or text input.

[1454] Input: User question (voice or text input)

[1455] Output: User questions saved on the device

[1456] Specific operation: The user asks a question using the device's microphone or keyboard, and the device saves the question.

[1457] Step 8:

[1458] Sending questions from the device to the server

[1459] The terminal sends the user's question to the server.

[1460] Input: User questions saved on the device

[1461] Output: The query data sent to the server

[1462] Specific operation: The terminal uses its wireless communication function to send the user's question to the server.

[1463] Step 9:

[1464] Re-explanation generation using generative AI

[1465] The server sends the received question to the generation AI, which then generates an explanation again.

[1466] Input: Question data

[1467] Output: Regenerated commentary

[1468] Specific operation: The generation AI analyzes the question content and generates a re-explanation if necessary.

[1469] Step 10:

[1470] Emotion recognition by server

[1471] The server sends the user's voice and video data to an emotion recognition engine to recognize the user's emotions.

[1472] Input: User audio and video data

[1473] Output: Emotion recognition result

[1474] Specific operation: The server inputs audio and video data into the emotion recognition engine and receives the analysis results.

[1475] Step 11:

[1476] Adjustment of commentary content

[1477] The generative AI adjusts the commentary content based on the emotion recognition results.

[1478] Input: Emotion recognition results and generated commentary

[1479] Output: Adjusted commentary

[1480] Specific operation: Match the emotion recognition results and reconstruct the commentary content to correspond to the user's emotions.

[1481] Step 12:

[1482] Sending adjustment results from the server to the device

[1483] The server sends the adjusted explanation received from the generation AI to the terminal.

[1484] Input: Adjusted commentary

[1485] Output: Adjustment results sent to the device

[1486] Specific operation: The server sends the adjusted commentary content to the terminal.

[1487] Step 13:

[1488] Displaying a recap on a terminal

[1489] The terminal displays the recap to the user.

[1490] Input: Adjustment description sent from the server

[1491] Output: A restatement that is displayed to the user

[1492] Specific operation: The terminal displays the adjusted commentary content on the screen for the user to see.

[1493] Through the above processing steps, the user can receive flexible learning support that is tailored to his or her level of understanding and emotions.

[1494] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1495] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1496] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1497] [Third embodiment]

[1498] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1499] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[1500] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1501] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1502] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1503] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1504] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1505] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1506] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1507] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1508] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1509] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[1510] This invention is a system that solves incomprehensible problems that arise when users study and supports deeper learning. This system provides a series of functions: the user inputs an image of a problem, and a generative AI automatically generates a solution and explanation, which are then output to the user. Furthermore, the system responds to follow-up questions from the user and provides further explanations to deepen the user's understanding.

[1511] System configuration

[1512] The system configuration is as follows:

[1513] 1. Terminal

[1514] The terminal is the device through which the user captures and inputs the image in question, and can include a smartphone, tablet, or PC.

[1515] 2. Server

[1516] The server receives and processes data sent from the device, analyzes the problem using image recognition technology, and sends the data to the generating AI.

[1517] 3. Generation AI

[1518] The AI ​​generator analyzes the received problem data and generates solutions and explanations. It also has the ability to adjust the explanations according to the user's level of understanding.

[1519] Program processing explanation

[1520] The processing of the program will be explained in natural language below.

[1521] ---

[1522] Problem capture and analysis

[1523] Terminal

[1524] The device prompts the user to take a picture of the problem, which is then saved on the device.

[1525] The device sends the saved image to the server.

[1526] server

[1527] The server receives the image sent from the terminal.

[1528] The server analyzes the received image and converts it into text data using image recognition technology (OCR).

[1529] The converted text data is formatted into a question and sent to the generation AI.

[1530] ---

[1531] Generate and display solutions and explanations

[1532] Generation AI

[1533] The generation AI analyzes the problem data it receives and generates a solution and explanation.

[1534] For example, the generated solution to the problem "2x + 3 = 7" would be the steps "2x = 7 - 3," "2x = 4," and "x = 2."

[1535] server

[1536] The server receives the solution and explanation from the generated AI.

[1537] The server sends the solution and explanation to the device.

[1538] Terminal

[1539] The terminal displays the solution and explanation to the user.

[1540] ---

[1541] Handling questions and restatements

[1542] User

[1543] If the user does not understand any part of the explanation, they can ask questions using voice recognition or text input.

[1544] Terminal

[1545] The device converts the user's speech into text.

[1546] The device sends the question to the server.

[1547] server

[1548] The server analyzes the received question and sends it to the generation AI.

[1549] Generation AI

[1550] Generative AI generates new explanations based on user questions.

[1551] server

[1552] The server receives the re-explanation from the generating AI and sends it to the terminal.

[1553] Terminal

[1554] The terminal displays the explanation again to the user.

[1555] If the user has further questions, the process is repeated.

[1556] ---

[1557] Specific examples

[1558] Example of a math problem

[1559] Consider the case where a user inputs the math problem "2x + 3 = 7." The following occurs:

[1560] 1. Problem capture

[1561] The user takes a picture of the problem using the device's camera.

[1562] The device sends the captured image to the server.

[1563] 2. Image Recognition and Problem Analysis

[1564] The server analyzes the received image and uses OCR to extract the text data "2x + 3 = 7".

[1565] The extracted text data is sent to the generation AI.

[1566] 3. Generating solutions and explanations

[1567] The generative AI generates the solutions to "2x + 3 = 7": "2x = 7 - 3," "2x = 4," and "x = 2," along with their explanations.

[1568] The server sends the generated solution and explanation to the terminal.

[1569] The terminal displays the solution and explanation to the user.

[1570] 4. Questions and Restatements

[1571] The user asks, "I don't understand why we subtract 3 from 7."

[1572] The device converts the speech into text and sends it to the server.

[1573] The server sends the question to the generation AI.

[1574] The generator AI generates a restatement: "To solve the equation, we need to shift the constant term, so we subtract 3 from 7."

[1575] The server sends the re-explanation to the terminal.

[1576] The terminal displays the re-explanation to the user.

[1577] In this way, questions can be asked and re-explained as many times as necessary until the user understands. This system allows users to learn at their own pace and achieve a deep understanding.

[1578] The processing flow will be explained below.

[1579] ---

[1580] Step 1:

[1581] The user takes a picture of the problem to be studied on the terminal.

[1582] The device will display instructions to the user saying, "Please take a picture of the problem."

[1583] The user takes a picture of the problem using the device's camera.

[1584] The device saves the captured image to its internal storage.

[1585] Step 2:

[1586] The device sends the saved image to the server.

[1587] The server receives the image data sent from the terminal.

[1588] The server temporarily stores the received image data.

[1589] Step 3:

[1590] The server uses image recognition technology (OCR) to extract text data from the image.

[1591] The server formats the extracted text data into a question format.

[1592] The server sends the formatted problem data to the generation AI.

[1593] Step 4:

[1594] The generation AI analyzes the problem data it receives.

[1595] The generative AI generates a solution to the provided problem.

[1596] Generative AI generates detailed explanations based on the solution.

[1597] Step 5:

[1598] The server receives the generated solution and explanation from the generated AI.

[1599] The server sends the received solution and explanation to the terminal.

[1600] Step 6:

[1601] The terminal displays the solution and explanation to the user.

[1602] The user checks the explanation and finds parts that he does not understand.

[1603] Step 7:

[1604] The user asks the terminal questions about anything they don't understand.

[1605] In the case of voice recognition, the user speaks a question into the terminal.

[1606] For text input: The user types the question into the device.

[1607] The device converts the user's speech into text (in the case of speech recognition).

[1608] The terminal sends a question from the user to the server.

[1609] Step 8:

[1610] The server receives the user's query.

[1611] The server sends the received question to the generation AI.

[1612] Step 9:

[1613] The generative AI generates a re-explanation based on the user's question.

[1614] The generating AI sends the new explanation to the server.

[1615] Step 10:

[1616] The server receives a re-explanation from the generated AI.

[1617] The server transmits the received explanation to the terminal.

[1618] Step 11:

[1619] The terminal displays the re-explanation to the user.

[1620] The user reviews the re-explanation and determines whether or not they understand it.

[1621] If not understood, the user asks the question again (return to step 7).

[1622] ---

[1623] By repeating this series of processes, the user can continue learning until their understanding deepens.

[1624] Example 1

[1625] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1626] In today's educational environment, users often encounter points they cannot understand while studying, and a system that can appropriately address these issues is needed. In particular, users who study independently have limited access to appropriate explanations and supplementary information. Therefore, the objective of this invention is to provide a system that can immediately resolve questions users encounter during the problem-solving process and deepen their understanding.

[1627] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1628] In this invention, the server includes a means for receiving an image of a problem provided by a user, an image recognition means for analyzing the problem from the received image, and a means having a generation AI for processing the analyzed problem data, which enables a means for a user to take a picture of the problem with a camera on a terminal and send it to the server, a means for the generation AI to analyze the problem based on the prompt and automatically generate steps for a solution and explanation, and a means for converting a user's voice input into text.

[1629] A "user" is an entity that uses the system to enter problems and receive solutions and explanations.

[1630] A "problem image" is a digital image that visually represents the problem the user wants to solve.

[1631] A "terminal" is a device that a user uses to capture and input the image in question, and includes smartphones, tablets, personal computers, etc.

[1632] A "server" is the central part of the system that receives and processes data sent from the terminals.

[1633] "Image recognition means" refers to techniques and devices that extract text data from received images.

[1634] "Text data" is character information extracted by image recognition means.

[1635] "Generative AI" is artificial intelligence that generates solutions and explanations based on analyzed problem data.

[1636] A "solution" is a procedure and computational process for solving a given problem.

[1637] "Explanation" is a text that explains each step of the solution and the background and reasons behind it.

[1638] A "prompt" is a form of text or data that is given to a generative AI as instructions or input to solve a problem.

[1639] "Voice input" refers to the act of inputting voice uttered by a user as information.

[1640] "Speech recognition means" refers to the technology and devices that convert a user's voice input into text.

[1641] A "re-explanation" is a new explanation generated by the AI ​​based on additional questions from the user.

[1642] This invention is a system that solves problems that arise when users are studying and supports deepening their learning. This system provides a series of functions in which the user inputs an image of a problem, and a generative AI automatically generates a solution and explanation, which are then output to the user. In addition, the system has a mechanism that responds to follow-up questions from the user and provides further explanations to deepen the user's understanding.

[1643] System configuration

[1644] 1. Terminal

[1645] The terminal is a device that users use to take and input images of the problem, and mainly includes smartphones, tablets, and PCs. It temporarily stores the images of the problem taken by the user and then sends them to the server.

[1646] 2. Server

[1647] The server receives the image sent from the device and uses Optical Character Recognition (OCR) technology to analyze the received image. For example, it uses a library such as Tesseract OCR to extract text data from the image. The extracted text data is then sent to the generation AI.

[1648] 3. Generation AI

[1649] The generating AI analyzes the problem according to the prompts based on the text data sent from the server, and generates a solution and explanation. The generating AI can use a generative model such as GPT. The generated solution and explanation are sent to the device via the server.

[1650] 4. Voice Recognition Methods

[1651] When a user has questions about something they don't understand, they can ask using voice recognition or text input. The device converts the user's voice into text. Voice recognition can use services such as the Google Speech-to-Text API or Amazon Transcribe.

[1652] Specific examples

[1653] Consider the case where a user inputs a math problem, "2x + 3 = 7." The system operates as follows:

[1654] 1. Problem capture

[1655] The user takes the image in question using the device's camera. At this time, the device launches the smartphone's camera app and displays a "Take a Photo" button. When the user presses the "Take a Photo" button, the image is captured and temporarily saved in the device's memory. The device then sends the image data to the server.

[1656] 2. Image Recognition and Problem Analysis

[1657] The server passes the received image data to a local image analysis engine, which uses OCR technology to extract the text data "2x + 3 = 7" from the image. This text data is then sent to the generation AI.

[1658] 3. Generating solutions and explanations

[1659] The generative AI analyzes the problem based on the prompt and generates the solutions "2x = 7 - 3", "2x = 4", "x = 2" and detailed step-by-step explanations. For example, the following prompts are used:

[1660] Please solve the equation 2x + 3 = 7 and explain each step.

[1661] The generated solution and explanation are sent to the terminal via the server.

[1662] 4. Displaying the solution and explanation

[1663] The device then displays the received solutions and explanations in a user-friendly format on the screen, with each step displayed step-by-step and accompanied by a detailed explanation for that step.

[1664] 5. Questions and Restatements

[1665] The user asks, "I don't understand why 3 is subtracted from 7." The device converts the user's voice into text using speech recognition technology and sends it to the server. The server sends the question as a prompt to the generation AI, which then generates a new explanation. This new explanation is sent to the device via the server and displayed to the user again.

[1666] In this way, questions and re-explanations can be repeated as many times as necessary until the user understands. Through this system, users can learn at their own pace and achieve a deep understanding.

[1667] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1668] Step 1: Take a picture of the problem and send it to us

[1669] Terminal

[1670] The user takes the image in question using the device's camera. For example, they launch the camera app on their smartphone and a "Take Photo" button appears.

[1671] When the user presses the "take a photo" button, an image is captured and temporarily saved in the device's internal memory.

[1672] The device sends the stored image data to the server using an HTTP POST request.

[1673] Input: A user-taken image of the problem

[1674] Output: Image data sent to the server

[1675] Specific behavior:

[1676] 1. The user activates the device camera.

[1677] 2. Press the "Capture" button to capture the image in question.

[1678] 3. Temporarily save image data.

[1679] 4. Send the image data to the server.

[1680] Step 2: Receiving and analyzing images

[1681] server

[1682] The server saves the image data received via the HTTP request in local storage.

[1683] The server passes the saved image data to an OCR (Optical Character Recognition) engine and begins processing.

[1684] OCR technology (e.g., Tesseract OCR) is used to extract text data from the image.

[1685] The extracted text data is formatted into questions and prompts are created to be sent to the generation AI.

[1686] Input: Image data sent from the device

[1687] Output: Problem text data to send to the generation AI

[1688] Specific behavior:

[1689] 1. The server receives the image data.

[1690] 2. Save the received image data locally.

[1691] 3. Pass the image data to the OCR engine.

[1692] 4. Extract text data using OCR technology.

[1693] 5. Format the text data into a question format.

[1694] 6. Create a prompt to send the formatted text data to the generation AI.

[1695] Step 3: Generate solutions and explanations

[1696] Generation AI

[1697] The generation AI receives the text data sent from the server.

[1698] The system analyzes the received text data as prompts and automatically generates steps to solve the problem and detailed explanations for each.

[1699] The generated solutions are in the form of, for example, "2x = 7 - 3," "2x = 4," or "x = 2." The explanation is something like, "First, subtract 3 on the left side from 7 on the right side."

[1700] Input: Question text data sent from the server

[1701] Output: Generated solution and explanation

[1702] Specific behavior:

[1703] 1. The generation AI receives the text data.

[1704] 2. Analyze text data as prompts.

[1705] 3. Automatically generate solutions and detailed explanations.

[1706] Step 4: Submit and view your solution and explanation

[1707] server

[1708] The server receives the generated solution and explanation from the generating AI.

[1709] The integrity of the received data is checked and it is sent to the terminal in JSON format or similar.

[1710] Input: Solution and explanation sent from the generating AI

[1711] Output: Solution and explanation data sent to the terminal

[1712] Specific behavior:

[1713] 1. The server receives the solution and explanation.

[1714] 2. Check data integrity.

[1715] 3. The data whose integrity has been confirmed is sent to the terminal.

[1716] Terminal

[1717] The terminal analyzes the solutions and explanations received from the server and displays them on the screen in a format that is easy for the user to see.

[1718] Each step is displayed step by step and a detailed explanation for each step is also provided.

[1719] Input: Solution and explanation data sent from the server

[1720] Output: The solution and explanation displayed to the user

[1721] Specific behavior:

[1722] 1. The device receives the solution and explanation.

[1723] 2. The solution and explanation are analyzed and displayed on the screen.

[1724] Step 5: Submit your question and voice input

[1725] User

[1726] If the user does not understand the explanation, they can enter a question using voice recognition or text input.

[1727] For example, tap the microphone icon on your smartphone to start voice input and say, "I don't know why I subtract 3 from 7."

[1728] Input: Additional questions about the explanation

[1729] Output: Input question data to the terminal

[1730] Specific behavior:

[1731] 1. The user taps the microphone icon.

[1732] 2. The user types in a question by voice.

[1733] Terminal

[1734] The device converts the user's speech into text using services such as the Google Speech-to-Text API or Amazon Transcribe.

[1735] The converted text data is sent to the server.

[1736] Input: A spoken question from the user

[1737] Output: Sends text-formatted question data to the server

[1738] Specific behavior:

[1739] 1. The device receives the audio data.

[1740] 2. Convert voice to text using voice recognition technology.

[1741] 3. Send the text data to the server.

[1742] Step 6: Parse the question and generate a restatement

[1743] server

[1744] The server receives the question sent from the terminal.

[1745] The received text data is sent to the generation AI as a prompt.

[1746] Input: Text question data sent from the terminal

[1747] Output: Prompt data to send to the generation AI

[1748] Specific behavior:

[1749] 1. The server receives the question text data.

[1750] 2. Send the text data to the generation AI as a prompt.

[1751] Generation AI

[1752] Generative AI generates new explanations based on user questions.

[1753] For example, it generates a restatement like "To solve the equation, we need to shift the constant term, so subtract 3 from 7."

[1754] Input: Question prompt sent from the server

[1755] Output: Generated restatement

[1756] Specific behavior:

[1757] 1. The generative AI receives a question prompt.

[1758] 2. Automatically generate new explanations based on questions.

[1759] Step 7: Submit and view the recap

[1760] server

[1761] The server receives the generated re-explanation from the generation AI and sends it to the terminal.

[1762] Input: Restatement sent by the generating AI

[1763] Output: Recap what is sent to the terminal

[1764] Specific behavior:

[1765] 1. The server receives the re-explanation.

[1766] 2. Send the re-explanation to the device.

[1767] Terminal

[1768] The terminal displays the re-explanation received from the server to the user.

[1769] Input: Recap sent by server

[1770] Output: A restatement that is displayed to the user

[1771] Specific behavior:

[1772] 1. The device receives the re-explanation.

[1773] 2. Display the re-explanation to the user.

[1774] In this way, the user can ask questions and receive re-explanations as many times as necessary until they understand. This system allows users to learn at their own pace and achieve a deep understanding.

[1775] (Application example 1)

[1776] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1777] In factories, work procedures are complex, and employees often make mistakes, leading to a decline in production efficiency and the inability to quickly resolve problems that arise during the work process. Newly introduced manuals and work procedures also tend to be poorly understood. A system is needed that allows employees to easily understand work procedures and receive assistance in resolving problems.

[1778] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1779] In this invention, the server includes means for receiving an image of a task provided by a user, image recognition means for analyzing the task from the received image, means having a generation AI for processing the analyzed task data, means for outputting to the user a solution and explanation obtained by the generation AI, means for receiving questions from the user, means for generating an explanation based on the received question and outputting it to the user, and means for displaying the generated solution and procedure through a smart device to support complex work procedures and problem solving in a factory. This enables employees to quickly solve problems they encounter during work and work efficiently.

[1780] The "means for receiving images of tasks provided by users" includes a device that allows a factory employee to take an image of a work procedure manual, a drawing, etc., and transmit the image to a server.

[1781] "Image recognition means" refers to technology that analyzes received images, identifies and extracts text and figures from the images, and converts them into data.

[1782] "Means having generative AI" refers to an artificial intelligence (AI) system that analyzes input data and automatically generates solutions and explanations.

[1783] "Means for outputting solutions and explanations to users" refers to technology for providing the solutions and explanations generated by the generation AI to factory employees via display devices or audio output devices.

[1784] The term "means for receiving questions" refers to technology that provides an interface for factory employees to input or voice questions and receives those questions.

[1785] "Means of generating new explanations based on questions and outputting them to users" refers to technology that uses a generative AI to reanalyze received questions, generate new explanations and procedures, and provide them to factory employees.

[1786] "Means for displaying generated solutions and procedures through smart devices to support complex work procedures and problem-solving within factories" refers to technology that uses wearable devices such as smart glasses or tablet terminals to display generated procedures and solutions to factory employees in real time.

[1787] The present invention is a system for supporting complex work procedures and problem solving within a factory, and in particular provides work procedures and solutions to factory employees using wearable devices such as smart glasses and tablet terminals.

[1788] System configuration

[1789] The system configuration is as follows:

[1790] 1. Terminal

[1791] Terminals are devices that factory employees use to take and input images of work procedures and drawings. These devices include smart glasses and tablets. These terminals are equipped with a camera function and send the images to a server.

[1792] 2. Server

[1793] The server is responsible for receiving and processing data sent from the terminal. The server has the following functions:

[1794] Image recognition means: The received image is analyzed and converted into text data using OCR (optical character recognition) technology. This converted data is sent to the generation AI.

[1795] Question analysis means: Receives and analyzes questions from factory employees. This data is also sent to the generation AI.

[1796] 3. Generation AI

[1797] The generative AI analyzes the received problem data and questions, and generates solutions and explanations. It also has the ability to adjust the explanation content according to the user's level of understanding. Examples of generative AI models used include GPT-4.

[1798] Program processing explanation

[1799] Problem capture and analysis

[1800] 1. The device instructs factory employees to take pictures of work procedures and drawings, which are then saved on the device.

[1801] 2. The device sends the saved image to the server.

[1802] 3. The server receives the image sent from the device.

[1803] 4. The server analyzes the received image and converts it into text data using OCR.

[1804] 5. The converted text data is formatted into a question and sent to the generation AI.

[1805] Generate and display solutions and explanations

[1806] 1. The generation AI analyzes the problem data it receives and generates specific work procedures and solutions.

[1807] 2. The generated procedures and solutions are sent to the terminal via the server.

[1808] 3. The terminal displays the generated solution and procedures to the factory worker.

[1809] Handling questions and restatements

[1810] 1. If a factory employee has questions about parts of the explanation that they don't understand, they can ask using voice recognition or text input.

[1811] 2. The device converts the user's speech into text.

[1812] 3. The device sends the question to the server.

[1813] 4. The server analyzes the received question and sends it to the generation AI.

[1814] 5. Generative AI generates new explanations based on the user's questions.

[1815] 6. The server receives the re-explanation from the generation AI and sends it to the device.

[1816] 7. The device displays the explanation again to the user.

[1817] Specific examples

[1818] Work procedure support example

[1819] Consider a case where a factory worker does not understand the work procedure "attach part A to part B." The following process is performed:

[1820] 1. A factory employee uses the device's camera to take an image of a work procedure manual or drawing.

[1821] 2. The device sends the captured image to the server.

[1822] 3. The server analyzes the received image and extracts text data using OCR.

[1823] 4. Send the extracted text data to the generation AI.

[1824] 5. The generative AI generates a solution and explanation such as "Steps for assembling parts A and B: 1. Remove the screws from part A. 2. Attach part A to part B."

[1825] 6. The generated procedure is sent to the terminal via the server and displayed to the factory employee.

[1826] 7. A factory worker asks, "Which way do you turn the screw?"

[1827] 8. The question is converted into text using voice recognition technology and sent to the server.

[1828] 9. The server sends the question to the generation AI, which generates a re-explanation: "Rotate clockwise."

[1829] 10. The re-explanation is sent to the terminal and displayed to the factory employee.

[1830] As a result, the system of the present invention can improve the work efficiency of users in a factory and speed up the problem solving process.

[1831] Prompt Sentence Examples

[1832] "Please solve the following problem: I don't understand step C: "Step C: Attach part A to part B.""

[1833] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1834] Step 1:

[1835] The terminal instructs the factory employee to take an image of the work procedure manual or blueprint. The factory employee takes an image of the work procedure manual or blueprint using the camera on the terminal. The input of this step is the factory employee's photographing behavior and the captured image, and the output is an image file saved on the terminal.

[1836] Step 2:

[1837] The device sends the saved image to the server. The device sends the image file taken by the factory employee to the server using an HTTP POST request. The input of this step is the saved image file, and the output is the image data received by the server.

[1838] Step 3:

[1839] The server receives the image sent from the terminal. The server temporarily stores the received image data. The input of this step is the image data sent from the terminal, and the output is the image data stored on the server.

[1840] Step 4:

[1841] The server analyzes the received image and converts it into text data using OCR. The server uses OCR technology (e.g., pytesseract) to extract text information from the image. The input for this step is the image data stored on the server, and the output is the extracted text data.

[1842] Step 5:

[1843] The server formats the converted text data into a question format and sends it to the generation AI. The server formats the text data into an appropriate format (e.g., JSON format) and sends an API request to the generation AI. The input of this step is the extracted text data, and the output is the question format data sent to the generation AI.

[1844] Step 6:

[1845] The generative AI analyzes the received problem data and generates a specific work procedure or solution. The generative AI (e.g., GPT-4) generates a solution and explanation based on the problem data. The input of this step is data in the form of a problem, and the output is the generated solution and explanation.

[1846] Step 7:

[1847] The server sends the generated solution and explanation to the terminal. The server sends the data received from the generation AI to the terminal. The input of this step is the generated solution and explanation, and the output is the solution and explanation sent to the terminal.

[1848] Step 8:

[1849] The terminal displays the generated solution and explanation to the factory employee. The terminal displays the received solution and explanation on the display. The input of this step is the solution and explanation sent to the terminal, and the output is display information that can be viewed by the factory employee.

[1850] Step 9:

[1851] If a factory employee does not understand a part of the explanation, they ask a question using voice recognition or text input. The factory employee then uses the terminal's voice input function or keyboard to input the question. The input for this step is the factory employee's voice or text input, and the output is the question data recognized by the terminal.

[1852] Step 10:

[1853] The terminal converts the user's voice into text. The terminal converts the voice data into text using voice recognition technology. The input of this step is the voice data of the factory employee, and the output is text data.

[1854] Step 11:

[1855] The terminal sends the question content to the server. The terminal then sends the question content converted into text data to the server. The input to this step is the text data, and the output is the question data sent to the server.

[1856] Step 12:

[1857] The server analyzes the received question and sends it to the generation AI. The server then formats the question and sends an API request to the generation AI again. The input of this step is the question data received by the server, and the output is the question data sent to the generation AI.

[1858] Step 13:

[1859] The generation AI generates a new explanation based on the user's question. The generation AI analyzes the question data and generates an appropriate re-explanation. The input of this step is the question data sent to the generation AI, and the output is the generated re-explanation data.

[1860] Step 14:

[1861] The server receives the re-explanation from the generation AI and sends it to the terminal. The server sends the generated re-explanation data to the terminal. The input of this step is the generated re-explanation data, and the output is the re-explanation data sent to the terminal.

[1862] Step 15:

[1863] The terminal displays the re-explanation to the factory employee. The terminal displays the re-explanation data on a display and provides it to the factory employee. The input of this step is the re-explanation data sent to the terminal, and the output is re-explanation information that can be viewed by the factory employee.

[1864] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1865] This invention is an educational system that not only solves problems that users do not understand when studying, but also recognizes the user's emotions and provides appropriate instruction methods. This system provides a series of functions: the user inputs an image of a problem, and a generative AI generates a solution and explanation, which is then output to the user. Furthermore, by incorporating a new emotion engine, the system has the function of recognizing the user's emotions and adjusting the explanation content according to the user's emotions.

[1866] System configuration

[1867] The system configuration is as follows:

[1868] 1. Terminal

[1869] The terminal is the device through which the user captures and inputs the image in question, and can include smartphones, tablets, and personal computers.

[1870] 2. Server

[1871] The server receives and processes data sent from the device, analyzes the problem using image recognition technology, and sends the data to the generating AI.

[1872] 3. Generation AI

[1873] The AI ​​generator analyzes the received problem data and generates solutions and explanations. It also has the ability to adjust the explanations according to the user's level of understanding.

[1874] 4. Emotion Engine

[1875] The emotion engine recognizes the user's emotions and provides a means for the system to respond based on those emotions. Specifically, it analyzes voice and facial expressions to determine whether the user understands or is confused.

[1876] Program processing explanation

[1877] The processing of the program will be explained in natural language below.

[1878] ---

[1879] Problem capture and analysis

[1880] Terminal

[1881] The device prompts the user to take a picture of the problem, which is then saved on the device.

[1882] The user takes a picture of the problem using the device's camera function.

[1883] The device sends the saved image to the server.

[1884] server

[1885] The server receives the image data sent from the terminal.

[1886] The server temporarily stores the received image data.

[1887] The server uses image recognition technology (OCR) to extract text data from the image.

[1888] The extracted text data is formatted into a question and sent to the generation AI.

[1889] ---

[1890] Generate and display solutions and explanations

[1891] Generation AI

[1892] The generation AI analyzes the problem data it receives.

[1893] The generative AI generates a solution to the provided problem.

[1894] Generative AI generates detailed explanations based on the solution.

[1895] server

[1896] The server receives the generated solution and explanation from the generated AI.

[1897] The server sends the received solution and explanation to the terminal.

[1898] Terminal

[1899] The terminal displays the solution and explanation to the user.

[1900] The user checks the explanation and finds parts that he does not understand.

[1901] ---

[1902] Handling questions and restatements

[1903] User

[1904] The user asks the terminal questions about anything they don't understand.

[1905] In the case of voice recognition, the user speaks a question into the terminal.

[1906] For text input: The user types the question into the device.

[1907] Terminal

[1908] The device converts the user's speech into text (in the case of speech recognition).

[1909] The terminal sends a question from the user to the server.

[1910] server

[1911] The server receives the user's query.

[1912] The server sends the received question to the generation AI.

[1913] Generation AI

[1914] The generative AI generates a re-explanation based on the user's question.

[1915] The generating AI sends the new explanation to the server.

[1916] server

[1917] The server receives a re-explanation from the generated AI.

[1918] The server transmits the received explanation to the terminal.

[1919] Terminal

[1920] The terminal displays the re-explanation to the user.

[1921] The user reviews the re-explanation and determines whether or not they understand it.

[1922] If the user does not understand, he or she asks the question again.

[1923] ---

[1924] Recognizing and Responding to Emotions

[1925] Terminal

[1926] The device collects the user's voice and camera footage.

[1927] The terminal transmits the collected data to the server.

[1928] server

[1929] The server sends the received data to the emotion engine.

[1930] The emotion engine analyzes voice and facial expressions to recognize the user's emotions.

[1931] The emotion engine sends the recognition results to the server.

[1932] Generation AI

[1933] The generative AI adjusts the commentary content based on the recognition results.

[1934] For example, if the user is confused, generate a more detailed explanation.

[1935] server

[1936] The server receives the adjusted commentary from the generated AI.

[1937] The server sends the adjusted commentary to the terminal.

[1938] Terminal

[1939] The terminal displays the adjusted description to the user.

[1940] ---

[1941] Specific examples

[1942] Example of a math problem

[1943] Consider the case where a user inputs the math problem "2x + 3 = 7." The following occurs:

[1944] 1. Problem capture

[1945] The user takes a picture of the problem using the device's camera.

[1946] The device sends the captured image to the server.

[1947] 2. Image Recognition and Problem Analysis

[1948] The server analyzes the received image and uses OCR to extract the text data "2x + 3 = 7".

[1949] The extracted text data is sent to the generation AI.

[1950] 3. Generating solutions and explanations

[1951] The generative AI generates the solutions to "2x + 3 = 7": "2x = 7 - 3," "2x = 4," and "x = 2," along with their explanations.

[1952] The server sends the generated solution and explanation to the terminal.

[1953] The terminal displays the solution and explanation to the user.

[1954] 4. Questions and Restatements

[1955] The user asks, "I don't understand why we subtract 3 from 7."

[1956] The device converts the speech into text and sends it to the server.

[1957] The server sends the question to the generation AI.

[1958] The generator AI generates a restatement: "To solve the equation, we need to shift the constant term, so we subtract 3 from 7."

[1959] The server sends the re-explanation to the terminal.

[1960] The terminal displays the re-explanation to the user.

[1961] 5. Emotion Recognition and Response

[1962] The device collects the user's voice and camera footage and sends them to the server.

[1963] The emotion engine analyzes voice and facial expressions to recognize the user's emotions.

[1964] The generative AI adjusts the commentary based on the perceived emotion (generating a more detailed explanation if the person is confused).

[1965] The server sends the adjusted commentary to the terminal, which displays it to the user.

[1966] In this way, users can receive support tailored to their level of understanding and emotional state. This system allows users to learn at their own pace and achieve deep understanding.

[1967] The processing flow will be explained below.

[1968] I understand. Now, I will explain the process flow of the invention that combines the emotion engine in concrete steps.

[1969] ---

[1970] Step 1:

[1971] The device will prompt the user to "Take a picture of the problem."

[1972] The user uses the device's camera to take the image in question.

[1973] The device saves the captured image to its internal storage.

[1974] Step 2:

[1975] The device sends the saved image to the server.

[1976] The server receives the image data sent from the terminal.

[1977] The server temporarily stores the received image data.

[1978] Step 3:

[1979] The server uses image recognition technology (OCR) to extract text data from the image.

[1980] The server formats the extracted text data into a question format.

[1981] The server sends the formatted problem data to the generation AI.

[1982] Step 4:

[1983] The generation AI analyzes the problem data it receives.

[1984] The generative AI generates a solution to the provided problem.

[1985] Generative AI generates detailed explanations based on the solution.

[1986] Step 5:

[1987] The server receives the generated solution and explanation from the generated AI.

[1988] The server sends the received solution and explanation to the terminal.

[1989] Step 6:

[1990] The terminal displays the solution and explanation to the user.

[1991] The user checks the explanation and finds parts that he does not understand.

[1992] Step 7:

[1993] The user asks the terminal questions about anything they don't understand.

[1994] In the case of voice recognition, the user speaks a question into the terminal.

[1995] For text input: The user types the question into the device.

[1996] Step 8:

[1997] The device converts the user's speech into text (in the case of speech recognition).

[1998] The terminal sends a question from the user to the server.

[1999] Step 9:

[2000] The server receives the user's query.

[2001] The server sends the received question to the generation AI.

[2002] Step 10:

[2003] The generative AI generates a re-explanation based on the user's question.

[2004] The generating AI sends the new explanation to the server.

[2005] Step 11:

[2006] The server receives a re-explanation from the generated AI.

[2007] The server transmits the received explanation to the terminal.

[2008] Step 12:

[2009] The terminal displays the re-explanation to the user.

[2010] The user reviews the re-explanation and determines whether or not they understand it.

[2011] If not understood, the user asks the question again (return to step 7).

[2012] Step 13:

[2013] The device collects the user's voice and camera footage.

[2014] The terminal transmits the collected data to the server.

[2015] Step 14:

[2016] The server sends the received data to the emotion engine.

[2017] The emotion engine analyzes voice and facial expressions to recognize the user's emotions.

[2018] The emotion engine sends the recognition results to the server.

[2019] Step 15:

[2020] The generative AI adjusts the commentary content based on the recognition results.

[2021] For example, if the user is confused, generate a more detailed explanation.

[2022] Step 16:

[2023] The server receives the adjusted commentary from the generated AI.

[2024] The server sends the adjusted commentary to the terminal.

[2025] Step 17:

[2026] The terminal displays the adjusted description to the user.

[2027] ---

[2028] Through this series of processes, users can receive support tailored to their level of understanding and emotional state.

[2029] Example 2

[2030] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[2031] In recent years, there has been a growing need for educational support systems that can quickly and accurately resolve user issues. However, conventional systems lack the explanations and suggestions needed to resolve issues that users cannot understand. Furthermore, because they do not take the user's emotional state into consideration, learning effectiveness may not be fully realized. To solve these problems, there is a need for the development of a system that not only provides answers to users' questions but also recognizes the user's emotions and provides appropriate guidance accordingly.

[2032] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[2033] In this invention, the server includes means for receiving images of problems provided by the user, image recognition means for analyzing the problems from the received images, artificial intelligence means for processing the analyzed problem data, means for outputting to the user the solutions and explanations obtained by the artificial intelligence means, means for receiving questions from the user, means for generating new explanations based on the received questions and outputting them to the user, emotion recognition means for collecting and analyzing emotion data of the user, and means for adjusting the content of the explanations based on the analysis results. This makes it possible to solve the user's understanding problems and provide appropriate guidance according to the user's emotional state, thereby maximizing the learning effect.

[2034] "User" refers to a person who uses the system.

[2035] "Question image" refers to an image file containing a study question that the user has taken using the camera function.

[2036] "Means for receiving" refers to the function that allows the system to receive problem images and questions provided by the user.

[2037] "Image recognition means" refers to technology (e.g., OCR technology) for extracting text data or problem information from received images.

[2038] "Generative AI" refers to a program that generates solutions and explanations from problem data analyzed using artificial intelligence technology.

[2039] "Means of output" refers to functions for providing users with the solutions and explanations obtained by the generating AI, such as screen display and audio output functions.

[2040] "Means for receiving questions" refers to a function for receiving additional questions or inquiries from users.

[2041] "Means for generating a new explanation" refers to the function of the generation AI to create a new explanation in response to a question from the user.

[2042] "Emotion data" refers to information about emotions analyzed from the user's voice, video, etc.

[2043] "Emotion recognition means" refers to technology for analyzing collected emotion data and recognizing the user's emotional state.

[2044] The "adjustment means" refers to a function for appropriately changing the commentary content based on the analysis results of the emotion recognition means.

[2045] This invention is an educational support system that not only resolves points that a user does not understand when studying, but also recognizes the user's emotions and provides appropriate teaching methods. This system is a combination of various hardware and software and consists of the following components:

[2046] First, the devices used by users include devices such as smartphones, tablets, and PCs. Users can use these devices to take pictures of study questions. The devices use their camera functions to acquire the images of the questions taken by the users.

[2047] The acquired image data is then sent to a server via a network. The server stores the received image data and uses image recognition technology (OCR = Optical Character Recognition) to extract text data from the image. This text data is then analyzed to determine the content of the problem.

[2048] The extracted text data is sent to the generative AI, which analyzes the problem based on the text data and generates a solution and a detailed explanation. The generative AI uses the latest machine learning and natural language processing technologies, enabling advanced analysis and generation.

[2049] The generated solutions and explanations are then sent to the terminal via the server. The terminal displays the received solutions and explanations to the user. The user can check the solutions and explanations displayed and ask additional questions about any parts they do not understand.

[2050] A user's question is entered into the device through text input or voice input. The device that receives the question sends the data to the server. The server then sends the user's question to the generation AI, which then generates a re-explanation. The generated re-explanation is sent to the device via the server and displayed to the user.

[2051] The system also incorporates emotion recognition functionality. The device collects the user's voice and camera footage and sends it to a server. The server then uses an emotion engine to analyze the user's emotion data and identify their emotional state. The emotion recognition data is then sent to a generation AI, which then adjusts the commentary content according to the user's emotional state. For example, if the user is confused, a more detailed explanation will be generated.

[2052] As a concrete example, consider the case where a user inputs the math problem "2x + 3 = 7." The user takes an image of the problem using the device's camera and sends it from the device to the server. The server uses OCR technology to extract the text data "2x + 3 = 7" and sends it to the generation AI. The generation AI generates a solution and detailed explanation for "2x + 3 = 7" and sends it to the device via the server. The device then displays the solution and explanation to the user.

[2053] If a user asks, "I don't know why we subtract 3 from 7," the device sends the question to the server, and the generation AI generates a restatement, saying, "Because we need to shift the constant term to solve the equation." The server sends the restatement to the device, which displays it to the user.

[2054] Here are some examples of specific prompts:

[2055] Problem: Generate a solution and explanation for "2x + 3 = 7".

[2056] User Question: "I don't understand why we subtract 3 from 7."

[2057] User sentiment: Confused

[2058] Please provide a more detailed explanation.

[2059] This system allows users to receive learning support at their own pace, and by receiving guidance tailored to their emotional state, they can achieve a deeper understanding.

[2060] The flow of the identification process in the second embodiment will be described with reference to FIG.

[2061] Step 1:

[2062] The device instructs the user to take a picture of the problem. Specifically, the device displays "Please take a picture of the problem." Input: User's instruction appears. Output: Message is displayed.

[2063] Step 2:

[2064] The user takes a picture of the problem using the device's camera. Input: The image of the problem taken by the user. Output: Image data.

[2065] Step 3:

[2066] The image captured by the device is sent to the server. Input: Image data from the device. Output: Image file sent to the server.

[2067] Step 4:

[2068] The server receives the image data. Input: Image data sent from the terminal. Output: Reception confirmation signal and storage of the image data.

[2069] Step 5:

[2070] The server uses image recognition technology (OCR) to extract text data. The server launches image analysis software to extract text data from the image. Input: Received image data. Output: Extracted text data.

[2071] Step 6:

[2072] The server sends the extracted text data to the generation AI. The server calls the generation AI's API and sends the text data. Input: Text data. Output: Data sent to the generation AI.

[2073] Step 7:

[2074] The generative AI generates a solution and explanation for the problem. The generative AI analyzes the input text data and generates an appropriate solution and detailed explanation. Input: Submitted text data. Output: Generated solution and explanation data.

[2075] Step 8:

[2076] The server sends the generated solution and explanation to the terminal. The server transfers the solution and explanation data received from the generating AI to the terminal. Input: Solution and explanation data from the generating AI. Output: Sending the solution and explanation data to the terminal.

[2077] Step 9:

[2078] The terminal displays the solution and explanation to the user. The solution and explanation are displayed on the terminal screen. Input: Solution and explanation data. Output: Display to user.

[2079] Step 10:

[2080] The user inputs a question. The user checks the solution and explanation and inputs a question about the part they don't understand. Voice input or text input is available. Input: User's question (text or voice). Output: Question data.

[2081] Step 11:

[2082] The terminal sends a question to the server. Input: Question data. Output: Question data sent to the server.

[2083] Step 12:

[2084] The server sends the question to the generation AI. The server calls the generation AI's API and sends the question data. Input: User's question data. Output: Sending question data to the generation AI.

[2085] Step 13:

[2086] The generation AI generates a restatement based on the question. The generation AI analyzes the question data and generates a restatement for the question. Input: Submitted question data. Output: Generated restatement data.

[2087] Step 14:

[2088] The server sends the re-explanation to the terminal. The server transfers the re-explanation data received from the generating AI to the terminal. Input: Re-explanation data from the generating AI. Output: Sending the re-explanation data to the terminal.

[2089] Step 15:

[2090] The terminal displays the re-explanation to the user. The re-explanation is displayed on the terminal screen. Input: Re-explanation data. Output: Re-display to user.

[2091] Step 16:

[2092] The device collects the user's emotional data and sends it to the server. The device collects the user's audio and video data and sends it to the server. Input: User's audio and video data. Output: Sending emotional data to the server.

[2093] Step 17:

[2094] The server sends emotion data to the emotion engine. Input: Emotion data sent from the device. Output: Data sent to the emotion engine.

[2095] Step 18:

[2096] The emotion engine analyzes the user's emotions. The emotion engine analyzes voice and facial expressions to recognize the user's emotional state. Input: Emotion data. Output: Analyzed emotional state data.

[2097] Step 19:

[2098] The emotion engine sends the analysis results to the generation AI. The emotion engine sends the analysis results to the generation AI, which adjusts the commentary content. Input: Analysis result data. Output: Data sent to the generation AI.

[2099] Step 20:

[2100] The generation AI adjusts the commentary content according to the emotion. The generation AI adjusts the commentary content based on the recognition results and generates an appropriate commentary. Input: Analysis result data. Output: Adjusted commentary data.

[2101] Step 21:

[2102] The server sends the adjusted commentary to the terminal. The server transfers the adjusted commentary data received from the generation AI to the terminal. Input: Adjusted commentary data from the generation AI. Output: Sending adjusted commentary data to the terminal.

[2103] Step 22:

[2104] The terminal displays the adjusted commentary to the user. The adjusted commentary is displayed on the terminal screen. Input: Adjusted commentary data. Output: Adjusted commentary displayed to the user.

[2105] (Application example 2)

[2106] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[2107] When a user encounters a point in the educational process that they do not understand, a common educational system has difficulty quickly recognizing that confusion and responding effectively. In particular, flexible responses based on the learner's emotions and level of understanding are required, but this has been difficult to achieve with conventional systems. In addition, there is a lack of means to grasp the learner's progress in real time and provide appropriate feedback on the spot.

[2108] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[2109] In this invention, the server includes means for receiving an image of a problem provided by a user, image recognition means for analyzing the problem from the received image, means having a generation AI for processing the analyzed problem data, means for outputting to the user a solution and explanation obtained by the generation AI, means for receiving a question from the user, means for generating an explanation based on the received question and outputting it to the user, means for identifying the user's emotion, and means for adjusting the content of the explanation based on the identified emotion. This enables flexible learning support according to the user's emotion and level of understanding, and also enables real-time progress confirmation and feedback.

[2110] "User" refers to the subject who uses this system to study, and who performs operations such as entering questions and asking questions.

[2111] "Means for receiving problem images" refers to an interface for receiving problem images photographed or scanned by a user and importing them into the system.

[2112] "Image recognition means" refers to technology that analyzes the received image of a question and extracts the question content, such as text data.

[2113] "Generative AI" refers to artificial intelligence technology that generates solutions and explanations based on extracted problem data.

[2114] "Means for outputting to the user" refers to an interface for visually or audibly conveying the solution and explanation obtained by the generative AI to the user.

[2115] The "means for receiving a question" refers to an interface for receiving a question within the system when a user inputs a question about a point that he or she does not understand.

[2116] The "means for generating a new explanation and outputting it to the user" refers to an interface for generating a new explanation based on a received user question and conveying that explanation to the user visually or audibly.

[2117] "Means for identifying user emotions" refers to technology that analyzes the user's facial expressions and voice to identify the emotion the user is currently feeling.

[2118] The "means for adjusting commentary content based on the identified emotion" refers to a technique for flexibly changing the generated commentary content according to the user's emotional state.

[2119] This invention is an educational system that not only solves problems that users do not understand when studying, but also recognizes the user's emotions and provides appropriate instruction methods. This system provides a series of functions: the user inputs an image of a problem, and a generative AI generates a solution and explanation, which is then output to the user. In addition, by incorporating a new emotion recognition engine, the system has the function of recognizing the user's emotions and adjusting the content of the explanation according to the user's emotions.

[2120] The main hardware and software used in implementing this invention are as follows:

[2121] Hardware

[2122] Smart Glasses (e.g., General Company Name, Smart Glasses)

[2123] server

[2124] Interface device (e.g., generic name, tablet, PC)

[2125] software

[2126] Emotion recognition engine (e.g., general company name, emotion recognition API)

[2127] Generative AI (e.g., general company name, AI generation engine)

[2128] Image recognition software using OCR technology

[2129] Voice Recognition Technology

[2130] System configuration and processing procedures

[2131] Device (smart glasses, etc.)

[2132] The user takes a picture of the problem and saves it on their device, which then sends the image data to the server.

[2133] server

[2134] The server receives the image data sent from the device, extracts text data from the image using OCR, and then formats the extracted text data into a question format and sends it to the generation AI.

[2135] Generation AI

[2136] The generation AI analyzes the received problem data, generates a solution and explanation, and sends the results back to the server.

[2137] server

[2138] The server receives the solution and explanation sent from the generated AI and sends it to the terminal, which displays the solution and explanation to the user.

[2139] User

[2140] The user checks the explanation and asks questions about parts they don't understand, and the device sends the questions to the server.

[2141] Re-explanation by generative AI

[2142] The server sends the question to the AI ​​generator, which then generates an explanation. The explanation is then sent to the device via the server and displayed to the user.

[2143] emotion recognition

[2144] The server receives the user's voice and video data and sends it to an emotion recognition engine. The emotion recognition engine analyzes the user's emotions and returns the results to the server. The server then sends the results to the generation AI, which adjusts the commentary content. The adjusted commentary is then sent to the device via the server and displayed to the user.

[2145] Specific examples

[2146] How to solve math problems

[2147] For example, a user inputs the math problem "2x + 3 = 7" into the system. The following specific prompts are fed into the generative AI model to get an explanation:

[2148] Explain in detail about "2x + 3 = 7," focusing on areas that students may find confusing. For example, explain in detail why the constant term is shifted.

[2149] In this way, the user can receive explanations in a way that is easy to understand. The system recognizes the user's emotions and adjusts the explanation content accordingly, providing effective learning support.

[2150] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[2151] Step 1:

[2152] Taking and receiving images with your device

[2153] The user takes a picture of the problem they want to learn using a device such as smart glasses, and the captured image is saved on the device.

[2154] Input: An image of the problem taken by the user

[2155] Output: Image data stored in the device

[2156] Specific operation: The user uses the device's camera function to take and save the image in question.

[2157] Step 2:

[2158] Sending images from the device to the server

[2159] The terminal transmits the captured image data to the server.

[2160] Input: Image data stored on the device

[2161] Output: Image data sent to the server

[2162] Specific operation: The terminal uses its wireless communication function to upload image data to the server.

[2163] Step 3:

[2164] Image analysis by server

[2165] The server analyzes the received image and extracts the text data using OCR technology.

[2166] Input: Image data sent to the server

[2167] Output: Extracted text data

[2168] Specific operation: Using OCR technology, the text information in the image is read and formatted as problem data.

[2169] Step 4:

[2170] Generative AI generates solutions and explanations

[2171] The server sends the extracted text data to the generation AI, which then generates a solution and explanation based on this.

[2172] Input: Text data sent from the server

[2173] Output: Generated solution and explanation

[2174] Specific operation: The generative AI analyzes problem data and generates the optimal solution and accompanying explanation.

[2175] Step 5:

[2176] Sending the generated results from the server to the device

[2177] The server sends the solution and explanation received from the generating AI to the terminal.

[2178] Input: Solutions and explanations generated by the generative AI

[2179] Output: Solution and explanation sent to terminal

[2180] Specific operation: The server receives the output results of the generation AI and sends them to the terminal.

[2181] Step 6:

[2182] Displaying solutions and explanations on your device

[2183] The terminal displays the received solution and explanation to the user.

[2184] Input: Solution and explanation sent from the server

[2185] Output: The solution and explanation displayed to the user

[2186] Specific operation: The terminal displays the solution and explanation on the screen for the user to see.

[2187] Step 7:

[2188] User Questions

[2189] Users can check the explanations and ask questions about anything they don't understand by voice or text input.

[2190] Input: User question (voice or text input)

[2191] Output: User questions saved on the device

[2192] Specific operation: The user asks a question using the device's microphone or keyboard, and the device saves the question.

[2193] Step 8:

[2194] Sending questions from the device to the server

[2195] The terminal sends the user's question to the server.

[2196] Input: User questions saved on the device

[2197] Output: The query data sent to the server

[2198] Specific operation: The terminal uses its wireless communication function to send the user's question to the server.

[2199] Step 9:

[2200] Re-explanation generation using generative AI

[2201] The server sends the received question to the generation AI, which then generates an explanation again.

[2202] Input: Question data

[2203] Output: Regenerated commentary

[2204] Specific operation: The generation AI analyzes the question content and generates a re-explanation if necessary.

[2205] Step 10:

[2206] Emotion recognition by server

[2207] The server sends the user's voice and video data to an emotion recognition engine to recognize the user's emotions.

[2208] Input: User audio and video data

[2209] Output: Emotion recognition result

[2210] Specific operation: The server inputs audio and video data into the emotion recognition engine and receives the analysis results.

[2211] Step 11:

[2212] Adjustment of commentary content

[2213] The generative AI adjusts the commentary content based on the emotion recognition results.

[2214] Input: Emotion recognition results and generated commentary

[2215] Output: Adjusted commentary

[2216] Specific operation: Match the emotion recognition results and reconstruct the commentary content to correspond to the user's emotions.

[2217] Step 12:

[2218] Sending adjustment results from the server to the device

[2219] The server sends the adjusted explanation received from the generation AI to the terminal.

[2220] Input: Adjusted commentary

[2221] Output: Adjustment results sent to the device

[2222] Specific operation: The server sends the adjusted commentary content to the terminal.

[2223] Step 13:

[2224] Displaying a recap on a terminal

[2225] The terminal displays the recap to the user.

[2226] Input: Adjustment description sent from the server

[2227] Output: A restatement that is displayed to the user

[2228] Specific operation: The terminal displays the adjusted commentary content on the screen for the user to see.

[2229] Through the above processing steps, the user can receive flexible learning support that is tailored to his or her level of understanding and emotions.

[2230] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[2231] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[2232] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[2233] [Fourth embodiment]

[2234] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[2235] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[2236] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[2237] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[2238] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[2239] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[2240] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[2241] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[2242] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[2243] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[2244] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[2245] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[2246] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2247] This invention is a system that solves incomprehensible problems that arise when users study and supports deeper learning. This system provides a series of functions: the user inputs an image of a problem, and a generative AI automatically generates a solution and explanation, which are then output to the user. Furthermore, the system responds to follow-up questions from the user and provides further explanations to deepen the user's understanding.

[2248] System configuration

[2249] The system configuration is as follows:

[2250] 1. Terminal

[2251] The terminal is the device through which the user captures and inputs the image in question, and can include a smartphone, tablet, or PC.

[2252] 2. Server

[2253] The server receives and processes data sent from the device, analyzes the problem using image recognition technology, and sends the data to the generating AI.

[2254] 3. Generation AI

[2255] The AI ​​generator analyzes the received problem data and generates solutions and explanations. It also has the ability to adjust the explanations according to the user's level of understanding.

[2256] Program processing explanation

[2257] The processing of the program will be explained in natural language below.

[2258] ---

[2259] Problem capture and analysis

[2260] Terminal

[2261] The device prompts the user to take a picture of the problem, which is then saved on the device.

[2262] The device sends the saved image to the server.

[2263] server

[2264] The server receives the image sent from the terminal.

[2265] The server analyzes the received image and converts it into text data using image recognition technology (OCR).

[2266] The converted text data is formatted into a question and sent to the generation AI.

[2267] ---

[2268] Generate and display solutions and explanations

[2269] Generation AI

[2270] The generation AI analyzes the problem data it receives and generates a solution and explanation.

[2271] For example, the generated solution to the problem "2x + 3 = 7" would be the steps "2x = 7 - 3," "2x = 4," and "x = 2."

[2272] server

[2273] The server receives the solution and explanation from the generated AI.

[2274] The server sends the solution and explanation to the device.

[2275] Terminal

[2276] The terminal displays the solution and explanation to the user.

[2277] ---

[2278] Handling questions and restatements

[2279] User

[2280] If the user does not understand any part of the explanation, they can ask questions using voice recognition or text input.

[2281] Terminal

[2282] The device converts the user's speech into text.

[2283] The device sends the question to the server.

[2284] server

[2285] The server analyzes the received question and sends it to the generation AI.

[2286] Generation AI

[2287] Generative AI generates new explanations based on user questions.

[2288] server

[2289] The server receives the re-explanation from the generating AI and sends it to the terminal.

[2290] Terminal

[2291] The terminal displays the explanation again to the user.

[2292] If the user has further questions, the process is repeated.

[2293] ---

[2294] Specific examples

[2295] Example of a math problem

[2296] Consider the case where a user inputs the math problem "2x + 3 = 7." The following occurs:

[2297] 1. Problem capture

[2298] The user takes a picture of the problem using the device's camera.

[2299] The device sends the captured image to the server.

[2300] 2. Image Recognition and Problem Analysis

[2301] The server analyzes the received image and uses OCR to extract the text data "2x + 3 = 7".

[2302] The extracted text data is sent to the generation AI.

[2303] 3. Generating solutions and explanations

[2304] The generative AI generates the solutions to "2x + 3 = 7": "2x = 7 - 3," "2x = 4," and "x = 2," along with their explanations.

[2305] The server sends the generated solution and explanation to the terminal.

[2306] The terminal displays the solution and explanation to the user.

[2307] 4. Questions and Restatements

[2308] The user asks, "I don't understand why we subtract 3 from 7."

[2309] The device converts the speech into text and sends it to the server.

[2310] The server sends the question to the generation AI.

[2311] The generator AI generates a restatement: "To solve the equation, we need to shift the constant term, so we subtract 3 from 7."

[2312] The server sends the re-explanation to the terminal.

[2313] The terminal displays the re-explanation to the user.

[2314] In this way, questions can be asked and re-explained as many times as necessary until the user understands. This system allows users to learn at their own pace and achieve a deep understanding.

[2315] The processing flow will be explained below.

[2316] ---

[2317] Step 1:

[2318] The user takes a picture of the problem to be studied on the terminal.

[2319] The device will display instructions to the user saying, "Please take a picture of the problem."

[2320] The user takes a picture of the problem using the device's camera.

[2321] The device saves the captured image to its internal storage.

[2322] Step 2:

[2323] The device sends the saved image to the server.

[2324] The server receives the image data sent from the terminal.

[2325] The server temporarily stores the received image data.

[2326] Step 3:

[2327] The server uses image recognition technology (OCR) to extract text data from the image.

[2328] The server formats the extracted text data into a question format.

[2329] The server sends the formatted problem data to the generation AI.

[2330] Step 4:

[2331] The generation AI analyzes the problem data it receives.

[2332] The generative AI generates a solution to the provided problem.

[2333] Generative AI generates detailed explanations based on the solution.

[2334] Step 5:

[2335] The server receives the generated solution and explanation from the generated AI.

[2336] The server sends the received solution and explanation to the terminal.

[2337] Step 6:

[2338] The terminal displays the solution and explanation to the user.

[2339] The user checks the explanation and finds parts that he does not understand.

[2340] Step 7:

[2341] The user asks the terminal questions about anything they don't understand.

[2342] In the case of voice recognition, the user speaks a question into the terminal.

[2343] For text input: The user types the question into the device.

[2344] The device converts the user's speech into text (in the case of speech recognition).

[2345] The terminal sends a question from the user to the server.

[2346] Step 8:

[2347] The server receives the user's query.

[2348] The server sends the received question to the generation AI.

[2349] Step 9:

[2350] The generative AI generates a re-explanation based on the user's question.

[2351] The generating AI sends the new explanation to the server.

[2352] Step 10:

[2353] The server receives a re-explanation from the generated AI.

[2354] The server transmits the received explanation to the terminal.

[2355] Step 11:

[2356] The terminal displays the re-explanation to the user.

[2357] The user reviews the re-explanation and determines whether or not they understand it.

[2358] If not understood, the user asks the question again (return to step 7).

[2359] ---

[2360] By repeating this series of processes, the user can continue learning until their understanding deepens.

[2361] Example 1

[2362] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2363] In today's educational environment, users often encounter points they cannot understand while studying, and a system that can appropriately address these issues is needed. In particular, users who study independently have limited access to appropriate explanations and supplementary information. Therefore, the objective of this invention is to provide a system that can immediately resolve questions users encounter during the problem-solving process and deepen their understanding.

[2364] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[2365] In this invention, the server includes a means for receiving an image of a problem provided by a user, an image recognition means for analyzing the problem from the received image, and a means having a generation AI for processing the analyzed problem data, which enables a means for a user to take a picture of the problem with a camera on a terminal and send it to the server, a means for the generation AI to analyze the problem based on the prompt and automatically generate steps for a solution and explanation, and a means for converting a user's voice input into text.

[2366] A "user" is an entity that uses the system to enter problems and receive solutions and explanations.

[2367] A "problem image" is a digital image that visually represents the problem the user wants to solve.

[2368] A "terminal" is a device that a user uses to capture and input the image in question, and includes smartphones, tablets, personal computers, etc.

[2369] A "server" is the central part of the system that receives and processes data sent from the terminals.

[2370] "Image recognition means" refers to techniques and devices that extract text data from received images.

[2371] "Text data" is character information extracted by image recognition means.

[2372] "Generative AI" is artificial intelligence that generates solutions and explanations based on analyzed problem data.

[2373] A "solution" is a procedure and computational process for solving a given problem.

[2374] "Explanation" is a text that explains each step of the solution and the background and reasons behind it.

[2375] A "prompt" is a form of text or data that is given to a generative AI as instructions or input to solve a problem.

[2376] "Voice input" refers to the act of inputting voice uttered by a user as information.

[2377] "Speech recognition means" refers to the technology and devices that convert a user's voice input into text.

[2378] A "re-explanation" is a new explanation generated by the AI ​​based on additional questions from the user.

[2379] This invention is a system that solves problems that arise when users are studying and supports deepening their learning. This system provides a series of functions in which the user inputs an image of a problem, and a generative AI automatically generates a solution and explanation, which are then output to the user. In addition, the system has a mechanism that responds to follow-up questions from the user and provides further explanations to deepen the user's understanding.

[2380] System configuration

[2381] 1. Terminal

[2382] The terminal is a device that users use to take and input images of the problem, and mainly includes smartphones, tablets, and PCs. It temporarily stores the images of the problem taken by the user and then sends them to the server.

[2383] 2. Server

[2384] The server receives the image sent from the device and uses Optical Character Recognition (OCR) technology to analyze the received image. For example, it uses a library such as Tesseract OCR to extract text data from the image. The extracted text data is then sent to the generation AI.

[2385] 3. Generation AI

[2386] The generating AI analyzes the problem according to the prompts based on the text data sent from the server, and generates a solution and explanation. The generating AI can use a generative model such as GPT. The generated solution and explanation are sent to the device via the server.

[2387] 4. Voice Recognition Methods

[2388] When a user has questions about something they don't understand, they can ask using voice recognition or text input. The device converts the user's voice into text. Voice recognition can use services such as the Google Speech-to-Text API or Amazon Transcribe.

[2389] Specific examples

[2390] Consider the case where a user inputs a math problem, "2x + 3 = 7." The system operates as follows:

[2391] 1. Problem capture

[2392] The user takes the image in question using the device's camera. At this time, the device launches the smartphone's camera app and displays a "Take a Photo" button. When the user presses the "Take a Photo" button, the image is captured and temporarily saved in the device's memory. The device then sends the image data to the server.

[2393] 2. Image Recognition and Problem Analysis

[2394] The server passes the received image data to a local image analysis engine, which uses OCR technology to extract the text data "2x + 3 = 7" from the image. This text data is then sent to the generation AI.

[2395] 3. Generating solutions and explanations

[2396] The generative AI analyzes the problem based on the prompt and generates the solutions "2x = 7 - 3", "2x = 4", "x = 2" and detailed step-by-step explanations. For example, the following prompts are used:

[2397] Please solve the equation 2x + 3 = 7 and explain each step.

[2398] The generated solution and explanation are sent to the terminal via the server.

[2399] 4. Displaying the solution and explanation

[2400] The device then displays the received solutions and explanations in a user-friendly format on the screen, with each step displayed step-by-step and accompanied by a detailed explanation for that step.

[2401] 5. Questions and Restatements

[2402] The user asks, "I don't understand why 3 is subtracted from 7." The device converts the user's voice into text using speech recognition technology and sends it to the server. The server sends the question as a prompt to the generation AI, which then generates a new explanation. This new explanation is sent to the device via the server and displayed to the user again.

[2403] In this way, questions and re-explanations can be repeated as many times as necessary until the user understands. Through this system, users can learn at their own pace and achieve a deep understanding.

[2404] The flow of the identification process in the first embodiment will be described with reference to FIG.

[2405] Step 1: Take a picture of the problem and send it to us

[2406] Terminal

[2407] The user takes the image in question using the device's camera. For example, they launch the camera app on their smartphone and a "Take Photo" button appears.

[2408] When the user presses the "take a photo" button, an image is captured and temporarily saved in the device's internal memory.

[2409] The device sends the stored image data to the server using an HTTP POST request.

[2410] Input: A user-taken image of the problem

[2411] Output: Image data sent to the server

[2412] Specific behavior:

[2413] 1. The user activates the device camera.

[2414] 2. Press the "Capture" button to capture the image in question.

[2415] 3. Temporarily save image data.

[2416] 4. Send the image data to the server.

[2417] Step 2: Receiving and analyzing images

[2418] server

[2419] The server saves the image data received via the HTTP request in local storage.

[2420] The server passes the saved image data to an OCR (Optical Character Recognition) engine and begins processing.

[2421] OCR technology (e.g., Tesseract OCR) is used to extract text data from the image.

[2422] The extracted text data is formatted into questions and prompts are created to be sent to the generation AI.

[2423] Input: Image data sent from the device

[2424] Output: Problem text data to send to the generation AI

[2425] Specific behavior:

[2426] 1. The server receives the image data.

[2427] 2. Save the received image data locally.

[2428] 3. Pass the image data to the OCR engine.

[2429] 4. Extract text data using OCR technology.

[2430] 5. Format the text data into a question format.

[2431] 6. Create a prompt to send the formatted text data to the generation AI.

[2432] Step 3: Generate solutions and explanations

[2433] Generation AI

[2434] The generation AI receives the text data sent from the server.

[2435] The system analyzes the received text data as prompts and automatically generates steps to solve the problem and detailed explanations for each.

[2436] The generated solutions are in the form of, for example, "2x = 7 - 3," "2x = 4," or "x = 2." The explanation is something like, "First, subtract 3 on the left side from 7 on the right side."

[2437] Input: Question text data sent from the server

[2438] Output: Generated solution and explanation

[2439] Specific behavior:

[2440] 1. The generation AI receives the text data.

[2441] 2. Analyze text data as prompts.

[2442] 3. Automatically generate solutions and detailed explanations.

[2443] Step 4: Submit and view your solution and explanation

[2444] server

[2445] The server receives the generated solution and explanation from the generating AI.

[2446] The integrity of the received data is checked and it is sent to the terminal in JSON format or similar.

[2447] Input: Solution and explanation sent from the generating AI

[2448] Output: Solution and explanation data sent to the terminal

[2449] Specific behavior:

[2450] 1. The server receives the solution and explanation.

[2451] 2. Check data integrity.

[2452] 3. The data whose integrity has been confirmed is sent to the terminal.

[2453] Terminal

[2454] The terminal analyzes the solutions and explanations received from the server and displays them on the screen in a format that is easy for the user to see.

[2455] Each step is displayed step by step and a detailed explanation for each step is also provided.

[2456] Input: Solution and explanation data sent from the server

[2457] Output: The solution and explanation displayed to the user

[2458] Specific behavior:

[2459] 1. The device receives the solution and explanation.

[2460] 2. The solution and explanation are analyzed and displayed on the screen.

[2461] Step 5: Submit your question and voice input

[2462] User

[2463] If the user does not understand the explanation, they can enter a question using voice recognition or text input.

[2464] For example, tap the microphone icon on your smartphone to start voice input and say, "I don't know why I subtract 3 from 7."

[2465] Input: Additional questions about the explanation

[2466] Output: Input question data to the terminal

[2467] Specific behavior:

[2468] 1. The user taps the microphone icon.

[2469] 2. The user types in a question by voice.

[2470] Terminal

[2471] The device converts the user's speech into text using services such as the Google Speech-to-Text API or Amazon Transcribe.

[2472] The converted text data is sent to the server.

[2473] Input: A spoken question from the user

[2474] Output: Sends text-formatted question data to the server

[2475] Specific behavior:

[2476] 1. The device receives the audio data.

[2477] 2. Convert voice to text using voice recognition technology.

[2478] 3. Send the text data to the server.

[2479] Step 6: Parse the question and generate a restatement

[2480] server

[2481] The server receives the question sent from the terminal.

[2482] The received text data is sent to the generation AI as a prompt.

[2483] Input: Text question data sent from the terminal

[2484] Output: Prompt data to send to the generation AI

[2485] Specific behavior:

[2486] 1. The server receives the question text data.

[2487] 2. Send the text data to the generation AI as a prompt.

[2488] Generation AI

[2489] Generative AI generates new explanations based on user questions.

[2490] For example, it generates a restatement like "To solve the equation, we need to shift the constant term, so subtract 3 from 7."

[2491] Input: Question prompt sent from the server

[2492] Output: Generated restatement

[2493] Specific behavior:

[2494] 1. The generative AI receives a question prompt.

[2495] 2. Automatically generate new explanations based on questions.

[2496] Step 7: Submit and view the recap

[2497] server

[2498] The server receives the generated re-explanation from the generation AI and sends it to the terminal.

[2499] Input: Restatement sent by the generating AI

[2500] Output: Recap what is sent to the terminal

[2501] Specific behavior:

[2502] 1. The server receives the re-explanation.

[2503] 2. Send the re-explanation to the device.

[2504] Terminal

[2505] The terminal displays the re-explanation received from the server to the user.

[2506] Input: Recap sent by server

[2507] Output: A restatement that is displayed to the user

[2508] Specific behavior:

[2509] 1. The device receives the re-explanation.

[2510] 2. Display the re-explanation to the user.

[2511] In this way, the user can ask questions and receive re-explanations as many times as necessary until they understand. This system allows users to learn at their own pace and achieve a deep understanding.

[2512] (Application example 1)

[2513] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2514] In factories, work procedures are complex, and employees often make mistakes, leading to a decline in production efficiency and the inability to quickly resolve problems that arise during the work process. Newly introduced manuals and work procedures also tend to be poorly understood. A system is needed that allows employees to easily understand work procedures and receive assistance in resolving problems.

[2515] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[2516] In this invention, the server includes means for receiving an image of a task provided by a user, image recognition means for analyzing the task from the received image, means having a generation AI for processing the analyzed task data, means for outputting to the user a solution and explanation obtained by the generation AI, means for receiving questions from the user, means for generating an explanation based on the received question and outputting it to the user, and means for displaying the generated solution and procedure through a smart device to support complex work procedures and problem solving in a factory. This enables employees to quickly solve problems they encounter during work and work efficiently.

[2517] The "means for receiving images of tasks provided by users" includes a device that allows a factory employee to take an image of a work procedure manual, a drawing, etc., and transmit the image to a server.

[2518] "Image recognition means" refers to technology that analyzes received images, identifies and extracts text and figures from the images, and converts them into data.

[2519] "Means having generative AI" refers to an artificial intelligence (AI) system that analyzes input data and automatically generates solutions and explanations.

[2520] "Means for outputting solutions and explanations to users" refers to technology for providing the solutions and explanations generated by the generation AI to factory employees via display devices or audio output devices.

[2521] The term "means for receiving questions" refers to technology that provides an interface for factory employees to input or voice questions and receives those questions.

[2522] "Means of generating new explanations based on questions and outputting them to users" refers to technology that uses a generative AI to reanalyze received questions, generate new explanations and procedures, and provide them to factory employees.

[2523] "Means for displaying generated solutions and procedures through smart devices to support complex work procedures and problem-solving within factories" refers to technology that uses wearable devices such as smart glasses or tablet terminals to display generated procedures and solutions to factory employees in real time.

[2524] The present invention is a system for supporting complex work procedures and problem solving within a factory, and in particular provides work procedures and solutions to factory employees using wearable devices such as smart glasses and tablet terminals.

[2525] System configuration

[2526] The system configuration is as follows:

[2527] 1. Terminal

[2528] Terminals are devices that factory employees use to take and input images of work procedures and drawings. These devices include smart glasses and tablets. These terminals are equipped with a camera function and send the images to a server.

[2529] 2. Server

[2530] The server is responsible for receiving and processing data sent from the terminal. The server has the following functions:

[2531] Image recognition means: The received image is analyzed and converted into text data using OCR (optical character recognition) technology. This converted data is sent to the generation AI.

[2532] Question analysis means: Receives and analyzes questions from factory employees. This data is also sent to the generation AI.

[2533] 3. Generation AI

[2534] The generative AI analyzes the received problem data and questions, and generates solutions and explanations. It also has the ability to adjust the explanation content according to the user's level of understanding. Examples of generative AI models used include GPT-4.

[2535] Program processing explanation

[2536] Problem capture and analysis

[2537] 1. The device instructs factory employees to take pictures of work procedures and drawings, which are then saved on the device.

[2538] 2. The device sends the saved image to the server.

[2539] 3. The server receives the image sent from the device.

[2540] 4. The server analyzes the received image and converts it into text data using OCR.

[2541] 5. The converted text data is formatted into a question and sent to the generation AI.

[2542] Generate and display solutions and explanations

[2543] 1. The generation AI analyzes the problem data it receives and generates specific work procedures and solutions.

[2544] 2. The generated procedures and solutions are sent to the terminal via the server.

[2545] 3. The terminal displays the generated solution and procedures to the factory worker.

[2546] Handling questions and restatements

[2547] 1. If a factory employee has questions about parts of the explanation that they don't understand, they can ask using voice recognition or text input.

[2548] 2. The device converts the user's speech into text.

[2549] 3. The device sends the question to the server.

[2550] 4. The server analyzes the received question and sends it to the generation AI.

[2551] 5. Generative AI generates new explanations based on the user's questions.

[2552] 6. The server receives the re-explanation from the generation AI and sends it to the device.

[2553] 7. The device displays the explanation again to the user.

[2554] Specific examples

[2555] Work procedure support example

[2556] Consider a case where a factory worker does not understand the work procedure "attach part A to part B." The following process is performed:

[2557] 1. A factory employee uses the device's camera to take an image of a work procedure manual or drawing.

[2558] 2. The device sends the captured image to the server.

[2559] 3. The server analyzes the received image and extracts text data using OCR.

[2560] 4. Send the extracted text data to the generation AI.

[2561] 5. The generative AI generates a solution and explanation such as "Steps for assembling parts A and B: 1. Remove the screws from part A. 2. Attach part A to part B."

[2562] 6. The generated procedure is sent to the terminal via the server and displayed to the factory employee.

[2563] 7. A factory worker asks, "Which way do you turn the screw?"

[2564] 8. The question is converted into text using voice recognition technology and sent to the server.

[2565] 9. The server sends the question to the generation AI, which generates a re-explanation: "Rotate clockwise."

[2566] 10. The re-explanation is sent to the terminal and displayed to the factory employee.

[2567] As a result, the system of the present invention can improve the work efficiency of users in a factory and speed up the problem solving process.

[2568] Prompt Sentence Examples

[2569] "Please solve the following problem: I don't understand step C: "Step C: Attach part A to part B.""

[2570] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[2571] Step 1:

[2572] The terminal instructs the factory employee to take an image of the work procedure manual or blueprint. The factory employee takes an image of the work procedure manual or blueprint using the camera on the terminal. The input of this step is the factory employee's photographing behavior and the captured image, and the output is an image file saved on the terminal.

[2573] Step 2:

[2574] The device sends the saved image to the server. The device sends the image file taken by the factory employee to the server using an HTTP POST request. The input of this step is the saved image file, and the output is the image data received by the server.

[2575] Step 3:

[2576] The server receives the image sent from the terminal. The server temporarily stores the received image data. The input of this step is the image data sent from the terminal, and the output is the image data stored on the server.

[2577] Step 4:

[2578] The server analyzes the received image and converts it into text data using OCR. The server uses OCR technology (e.g., pytesseract) to extract text information from the image. The input for this step is the image data stored on the server, and the output is the extracted text data.

[2579] Step 5:

[2580] The server formats the converted text data into a question format and sends it to the generation AI. The server formats the text data into an appropriate format (e.g., JSON format) and sends an API request to the generation AI. The input of this step is the extracted text data, and the output is the question format data sent to the generation AI.

[2581] Step 6:

[2582] The generative AI analyzes the received problem data and generates a specific work procedure or solution. The generative AI (e.g., GPT-4) generates a solution and explanation based on the problem data. The input of this step is data in the form of a problem, and the output is the generated solution and explanation.

[2583] Step 7:

[2584] The server sends the generated solution and explanation to the terminal. The server sends the data received from the generation AI to the terminal. The input of this step is the generated solution and explanation, and the output is the solution and explanation sent to the terminal.

[2585] Step 8:

[2586] The terminal displays the generated solution and explanation to the factory employee. The terminal displays the received solution and explanation on the display. The input of this step is the solution and explanation sent to the terminal, and the output is display information that can be viewed by the factory employee.

[2587] Step 9:

[2588] If a factory employee does not understand a part of the explanation, they ask a question using voice recognition or text input. The factory employee then uses the terminal's voice input function or keyboard to input the question. The input for this step is the factory employee's voice or text input, and the output is the question data recognized by the terminal.

[2589] Step 10:

[2590] The terminal converts the user's voice into text. The terminal converts the voice data into text using voice recognition technology. The input of this step is the voice data of the factory employee, and the output is text data.

[2591] Step 11:

[2592] The terminal sends the question content to the server. The terminal then sends the question content converted into text data to the server. The input to this step is the text data, and the output is the question data sent to the server.

[2593] Step 12:

[2594] The server analyzes the received question and sends it to the generation AI. The server then formats the question and sends an API request to the generation AI again. The input of this step is the question data received by the server, and the output is the question data sent to the generation AI.

[2595] Step 13:

[2596] The generation AI generates a new explanation based on the user's question. The generation AI analyzes the question data and generates an appropriate re-explanation. The input of this step is the question data sent to the generation AI, and the output is the generated re-explanation data.

[2597] Step 14:

[2598] The server receives the re-explanation from the generation AI and sends it to the terminal. The server sends the generated re-explanation data to the terminal. The input of this step is the generated re-explanation data, and the output is the re-explanation data sent to the terminal.

[2599] Step 15:

[2600] The terminal displays the re-explanation to the factory employee. The terminal displays the re-explanation data on a display and provides it to the factory employee. The input of this step is the re-explanation data sent to the terminal, and the output is re-explanation information that can be viewed by the factory employee.

[2601] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[2602] This invention is an educational system that not only solves problems that users do not understand when studying, but also recognizes the user's emotions and provides appropriate instruction methods. This system provides a series of functions: the user inputs an image of a problem, and a generative AI generates a solution and explanation, which is then output to the user. Furthermore, by incorporating a new emotion engine, the system has the function of recognizing the user's emotions and adjusting the explanation content according to the user's emotions.

[2603] System configuration

[2604] The system configuration is as follows:

[2605] 1. Terminal

[2606] The terminal is the device through which the user captures and inputs the image in question, and can include smartphones, tablets, and personal computers.

[2607] 2. Server

[2608] The server receives and processes data sent from the device, analyzes the problem using image recognition technology, and sends the data to the generating AI.

[2609] 3. Generation AI

[2610] The AI ​​generator analyzes the received problem data and generates solutions and explanations. It also has the ability to adjust the explanations according to the user's level of understanding.

[2611] 4. Emotion Engine

[2612] The emotion engine recognizes the user's emotions and provides a means for the system to respond based on those emotions. Specifically, it analyzes voice and facial expressions to determine whether the user understands or is confused.

[2613] Program processing explanation

[2614] The processing of the program will be explained in natural language below.

[2615] ---

[2616] Problem capture and analysis

[2617] Terminal

[2618] The device prompts the user to take a picture of the problem, which is then saved on the device.

[2619] The user takes a picture of the problem using the device's camera function.

[2620] The device sends the saved image to the server.

[2621] server

[2622] The server receives the image data sent from the terminal.

[2623] The server temporarily stores the received image data.

[2624] The server uses image recognition technology (OCR) to extract text data from the image.

[2625] The extracted text data is formatted into a question and sent to the generation AI.

[2626] ---

[2627] Generate and display solutions and explanations

[2628] Generation AI

[2629] The generation AI analyzes the problem data it receives.

[2630] The generative AI generates a solution to the provided problem.

[2631] Generative AI generates detailed explanations based on the solution.

[2632] server

[2633] The server receives the generated solution and explanation from the generated AI.

[2634] The server sends the received solution and explanation to the terminal.

[2635] Terminal

[2636] The terminal displays the solution and explanation to the user.

[2637] The user checks the explanation and finds parts that he does not understand.

[2638] ---

[2639] Handling questions and restatements

[2640] User

[2641] The user asks the terminal questions about anything they don't understand.

[2642] In the case of voice recognition, the user speaks a question into the terminal.

[2643] For text input: The user types the question into the device.

[2644] Terminal

[2645] The device converts the user's speech into text (in the case of speech recognition).

[2646] The terminal sends a question from the user to the server.

[2647] server

[2648] The server receives the user's query.

[2649] The server sends the received question to the generation AI.

[2650] Generation AI

[2651] The generative AI generates a re-explanation based on the user's question.

[2652] The generating AI sends the new explanation to the server.

[2653] server

[2654] The server receives a re-explanation from the generated AI.

[2655] The server transmits the received explanation to the terminal.

[2656] Terminal

[2657] The terminal displays the re-explanation to the user.

[2658] The user reviews the re-explanation and determines whether or not they understand it.

[2659] If the user does not understand, he or she asks the question again.

[2660] ---

[2661] Recognizing and Responding to Emotions

[2662] Terminal

[2663] The device collects the user's voice and camera footage.

[2664] The terminal transmits the collected data to the server.

[2665] server

[2666] The server sends the received data to the emotion engine.

[2667] The emotion engine analyzes voice and facial expressions to recognize the user's emotions.

[2668] The emotion engine sends the recognition results to the server.

[2669] Generation AI

[2670] The generative AI adjusts the commentary content based on the recognition results.

[2671] For example, if the user is confused, generate a more detailed explanation.

[2672] server

[2673] The server receives the adjusted commentary from the generated AI.

[2674] The server sends the adjusted commentary to the terminal.

[2675] Terminal

[2676] The terminal displays the adjusted description to the user.

[2677] ---

[2678] Specific examples

[2679] Example of a math problem

[2680] Consider the case where a user inputs the math problem "2x + 3 = 7." The following occurs:

[2681] 1. Problem capture

[2682] The user takes a picture of the problem using the device's camera.

[2683] The device sends the captured image to the server.

[2684] 2. Image Recognition and Problem Analysis

[2685] The server analyzes the received image and uses OCR to extract the text data "2x + 3 = 7".

[2686] The extracted text data is sent to the generation AI.

[2687] 3. Generating solutions and explanations

[2688] The generative AI generates the solutions to "2x + 3 = 7": "2x = 7 - 3," "2x = 4," and "x = 2," along with their explanations.

[2689] The server sends the generated solution and explanation to the terminal.

[2690] The terminal displays the solution and explanation to the user.

[2691] 4. Questions and Restatements

[2692] The user asks, "I don't understand why we subtract 3 from 7."

[2693] The device converts the speech into text and sends it to the server.

[2694] The server sends the question to the generation AI.

[2695] The generator AI generates a restatement: "To solve the equation, we need to shift the constant term, so we subtract 3 from 7."

[2696] The server sends the re-explanation to the terminal.

[2697] The terminal displays the re-explanation to the user.

[2698] 5. Emotion Recognition and Response

[2699] The device collects the user's voice and camera footage and sends them to the server.

[2700] The emotion engine analyzes voice and facial expressions to recognize the user's emotions.

[2701] The generative AI adjusts the commentary based on the perceived emotion (generating a more detailed explanation if the person is confused).

[2702] The server sends the adjusted commentary to the terminal, which displays it to the user.

[2703] In this way, users can receive support tailored to their level of understanding and emotional state. This system allows users to learn at their own pace and achieve deep understanding.

[2704] The processing flow will be explained below.

[2705] I understand. Now, I will explain the process flow of the invention that combines the emotion engine in concrete steps.

[2706] ---

[2707] Step 1:

[2708] The device will prompt the user to "Take a picture of the problem."

[2709] The user uses the device's camera to take the image in question.

[2710] The device saves the captured image to its internal storage.

[2711] Step 2:

[2712] The device sends the saved image to the server.

[2713] The server receives the image data sent from the terminal.

[2714] The server temporarily stores the received image data.

[2715] Step 3:

[2716] The server uses image recognition technology (OCR) to extract text data from the image.

[2717] The server formats the extracted text data into a question format.

[2718] The server sends the formatted problem data to the generation AI.

[2719] Step 4:

[2720] The generation AI analyzes the problem data it receives.

[2721] The generative AI generates a solution to the provided problem.

[2722] Generative AI generates detailed explanations based on the solution.

[2723] Step 5:

[2724] The server receives the generated solution and explanation from the generated AI.

[2725] The server sends the received solution and explanation to the terminal.

[2726] Step 6:

[2727] The terminal displays the solution and explanation to the user.

[2728] The user checks the explanation and finds parts that he does not understand.

[2729] Step 7:

[2730] The user asks the terminal questions about anything they don't understand.

[2731] In the case of voice recognition, the user speaks a question into the terminal.

[2732] For text input: The user types the question into the device.

[2733] Step 8:

[2734] The device converts the user's speech into text (in the case of speech recognition).

[2735] The terminal sends a question from the user to the server.

[2736] Step 9:

[2737] The server receives the user's query.

[2738] The server sends the received question to the generation AI.

[2739] Step 10:

[2740] The generative AI generates a re-explanation based on the user's question.

[2741] The generating AI sends the new explanation to the server.

[2742] Step 11:

[2743] The server receives a re-explanation from the generated AI.

[2744] The server transmits the received explanation to the terminal.

[2745] Step 12:

[2746] The terminal displays the re-explanation to the user.

[2747] The user reviews the re-explanation and determines whether or not they understand it.

[2748] If not understood, the user asks the question again (return to step 7).

[2749] Step 13:

[2750] The device collects the user's voice and camera footage.

[2751] The terminal transmits the collected data to the server.

[2752] Step 14:

[2753] The server sends the received data to the emotion engine.

[2754] The emotion engine analyzes voice and facial expressions to recognize the user's emotions.

[2755] The emotion engine sends the recognition results to the server.

[2756] Step 15:

[2757] The generative AI adjusts the commentary content based on the recognition results.

[2758] For example, if the user is confused, generate a more detailed explanation.

[2759] Step 16:

[2760] The server receives the adjusted commentary from the generated AI.

[2761] The server sends the adjusted commentary to the terminal.

[2762] Step 17:

[2763] The terminal displays the adjusted description to the user.

[2764] ---

[2765] Through this series of processes, users can receive support tailored to their level of understanding and emotional state.

[2766] Example 2

[2767] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2768] In recent years, there has been a growing need for educational support systems that can quickly and accurately resolve user issues. However, conventional systems lack the explanations and suggestions needed to resolve issues that users cannot understand. Furthermore, because they do not take the user's emotional state into consideration, learning effectiveness may not be fully realized. To solve these problems, there is a need for the development of a system that not only provides answers to users' questions but also recognizes the user's emotions and provides appropriate guidance accordingly.

[2769] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[2770] In this invention, the server includes means for receiving images of problems provided by the user, image recognition means for analyzing the problems from the received images, artificial intelligence means for processing the analyzed problem data, means for outputting to the user the solutions and explanations obtained by the artificial intelligence means, means for receiving questions from the user, means for generating new explanations based on the received questions and outputting them to the user, emotion recognition means for collecting and analyzing emotion data of the user, and means for adjusting the content of the explanations based on the analysis results. This makes it possible to solve the user's understanding problems and provide appropriate guidance according to the user's emotional state, thereby maximizing the learning effect.

[2771] "User" refers to a person who uses the system.

[2772] "Question image" refers to an image file containing a study question that the user has taken using the camera function.

[2773] "Means for receiving" refers to the function that allows the system to receive problem images and questions provided by the user.

[2774] "Image recognition means" refers to technology (e.g., OCR technology) for extracting text data or problem information from received images.

[2775] "Generative AI" refers to a program that generates solutions and explanations from problem data analyzed using artificial intelligence technology.

[2776] "Means of output" refers to functions for providing users with the solutions and explanations obtained by the generating AI, such as screen display and audio output functions.

[2777] "Means for receiving questions" refers to a function for receiving additional questions or inquiries from users.

[2778] "Means for ...

Claims

1. means for receiving a user-provided problem image; image recognition means for analyzing a problem from the received image; means having a generating AI for processing the analyzed problem data; A means for outputting the solution and explanation obtained by the generating AI to the user; means for receiving a query from a user; a means for generating an explanation based on the received question and outputting the explanation to the user; A system including:

2. 10. The system of claim 1, further comprising speech recognition means for converting said user-provided speech into text.

3. The system according to claim 1, wherein the generating AI includes means for adjusting the content of the explanation according to the user's level of understanding.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A